2606.20560v1 Jun 18, 2026 cs.LG

DiffusionGemma의 투명성은 어느 정도인가?

How Transparent is DiffusionGemma?

Arthur Conmy
Arthur Conmy
Citations: 38
h-index: 4
Neel Nanda
Neel Nanda
Citations: 11,588
h-index: 37
Joshua Engels
Joshua Engels
Citations: 447
h-index: 9
Bilal Chughtai
Bilal Chughtai
Citations: 593
h-index: 11
Rohin Shah
Rohin Shah
Citations: 921
h-index: 9
C. McDougall
C. McDougall
Citations: 735
h-index: 4
János Kramár
János Kramár
Citations: 298
h-index: 3
Senthoran Rajamanoharan
Senthoran Rajamanoharan
Citations: 0
h-index: 0
Cindy Wu
Cindy Wu
Citations: 168
h-index: 3
Asic Chen
Asic Chen
Citations: 12
h-index: 1
Jean Tarbouriech
Jean Tarbouriech
Citations: 3,563
h-index: 11
Mingdao Ma
Mingdao Ma
Citations: 0
h-index: 0
Brendan O'Donoghue
Brendan O'Donoghue
Citations: 8,069
h-index: 26
Joao Gabriel Lopes de Oliveira
Joao Gabriel Lopes de Oliveira
Citations: 0
h-index: 0

LLM 추론 과정의 투명성은 모델 결정에 대한 이해를 높이고, 오용 및 오류를 줄이며, 예상치 못한 모델 동작을 디버깅하는 데 중요한 역할을 합니다. 하지만 DiffusionGemma는 연산의 상당 부분을 연속적인 잠재 공간에서 수행합니다. 이로 인해 DiffusionGemma의 추론 과정이 덜 투명해질까요? 우리는 투명성을 변수 투명성과 알고리즘 투명성이라는 두 가지 구성 요소로 나누어 이 질문을 연구했습니다. 변수 투명성은 모델의 연산 상태에 대한 중간 단계를 이해하는 정도를 의미하며, 알고리즘 투명성은 이러한 단계를 사용하여 모델이 어떻게 결과를 도출했는지 재구성할 수 있는 능력을 의미합니다. 초기 분석 결과, DiffusionGemma는 변수 투명성이 낮았습니다. 해석 가능한 모델 상태 사이에서 발생하는 '불투명한 연산 깊이'가 해당되는 오토리거시브 Gemma 4 모델보다 약 28.6배 더 높게 나타났습니다. 하지만 우리는 해석 가능한 토큰 병목 현상을 통해 노이즈 제거 단계 간의 정보 흐름을 매핑함으로써, 하위 작업 성능 저하 없이 불투명한 연산 깊이를 Gemma 4 모델의 1.1배 수준으로 줄일 수 있음을 보여줍니다. 이러한 중간 상태를 해석 가능하게 취급하면 불투명한 연산 깊이가 Gemma 4 모델보다 훨씬 낮아집니다. 확산 모델은 오토리거시브 모델에 비해 알고리즘 투명성을 확보하기 어렵습니다. 그 이유는 캔버스 내의 모든 토큰 예측이 매 노이즈 제거 단계마다 변경될 수 있기 때문이며, 이는 모델이 노이즈 제거 과정에서 복잡한 분산 알고리즘을 구현할 수 있는 능력을 부여합니다. 이러한 격차를 줄이기 위해 우리는 다양한 해석 가능성 연구 사례를 수행하여 비선형적 추론, 토큰 및 시퀀스 스머징, 중간 컨텍스트 추론과 같은 확산 모델에 특유한 현상에 대한 초기 증거를 발견했습니다. 마지막으로, 투명성의 핵심 응용 분야인 '모니터링 가능성'을 테스트하여 모델 출력 결과가 하위 작업에 얼마나 유용한지를 측정했습니다. 그 결과, DiffusionGemma는 Gemma 4와 유사한 수준의 모니터링 가능성을 보이는 것으로 나타났습니다.

Original Abstract

LLM reasoning transparency is a critical affordance for understanding model decisions, mitigating misuse and misalignment, and debugging surprising model behaviors. However, DiffusionGemma performs a larger fraction of its computation in a continuous latent space; does this make its reasoning less transparent? We study this question by decomposing transparency into two components: variable transparency, whether we understand intermediate snapshots of a model's computational state; and algorithmic transparency, whether we can use these snapshots to reconstruct the process by which the model arrived at its outputs. Naively, DiffusionGemma has poor variable transparency: its opaque serial depth, the amount of serial computation that occurs in between interpretable model states, seems at first 28.6X higher than the corresponding autoregressive Gemma 4 model. However, we show that we can map the information flowing between denoising steps through an interpretable token bottleneck with no decrease in downstream performance. Treating these intermediate states as interpretable reduces the opaque serial depth to just 1.1X that of Gemma 4. Algorithmic transparency is harder for diffusion models than for autoregressive models because all token predictions in the canvas can change at every denoising step, giving the model the power to implement complicated distributed algorithms during the denoising process. To begin bridging this gap, we conduct a suite of interpretability case studies, uncovering initial evidence of novel diffusion-specific phenomena such as non-chronological reasoning, token and sequence smearing, and intermediate-context reasoning. Finally, we test monitorability, a key application of transparency that measures whether model outputs are useful for downstream tasks. We find that DiffusionGemma is similarly monitorable to Gemma 4.

0 Citations
0 Influential
18.5 Altmetric
92.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!