누가 디코딩을 이끌어야 하는가? 마스크된 확산 언어 모델의 앙상블을 위한 신뢰할 수 있는 경로 추적
Who Should Lead Decoding Now? Tracking Reliable Trajectories for Ensembling Masked Diffusion Language Models
마스크된 확산 언어 모델(MDLM)은 시퀀스 생성에 대한 새로운 패러다임으로 등장했습니다. MDLM이 기능과 지식 범위 측면에서 다양해짐에 따라, 이들의 지식을 어떻게 결합할 것인가 하는 중요한 질문이 제기됩니다. 이에 앞서, 우리는 MDLM의 독특한 디코딩 동역학을 조사합니다. 성공적인 생성 결과는 답변 관련 위치에서 안정적인 신뢰도 변화를 보이는 반면, 신뢰성이 낮은 경로는 종종 다른 모델로부터 유망한 중간 상태를 주입하여 수정될 수 있습니다. 이러한 관찰에 따라, 우리는 $ extbf{TIE}$ ($ extbf{T}$rajectory-based $ extbf{I}$terative $ extbf{E}$nsembling)라는 지식 융합 프레임워크를 제안합니다. TIE는 MDLM이 반복적으로 신뢰할 수 있는 디코딩 경로를 식별하고 모델 간에 이를 전달합니다. TIE는 답변 관련 위치에서의 신뢰도 변화를 추적하여 현재 더 신뢰할 수 있는 경로를 따르는 모델을 결정하고, 부분적으로 노이즈 제거된 시퀀스를 선택적으로 다른 모델로 전송합니다. 더 유망한 경로를 따르는 모델은 종종 디노이징 단계에 따라 변경되므로, TIE는 다양한 모델이 생성 과정의 서로 다른 단계에서 상호 보완적인 강점을 기여할 수 있도록 합니다. 다양한 추론 작업에서의 강력한 성능과 함께, 우리의 분석 결과는 TIE가 MDLM 앙상블이라는 아직 탐구되지 않은 문제에 대한 실용적인 접근 방식을 제공한다는 것을 시사합니다.
Masked Diffusion Language Models (MDLMs) have emerged as a distinct paradigm for sequence generation. As MDLMs become diverse in capabilities and knowledge coverage, an important question is how to combine their knowledge. Toward this, we first investigate the unique decoding dynamics of MDLMs. We find that successful generations exhibit stable confidence dynamics over answer-relevant positions, while unreliable trajectories can often be corrected by injecting promising intermediate states from other models. Guided by this observation, we propose $\textbf{TIE}$ ($\textbf{T}$rajectory-based $\textbf{I}$terative $\textbf{E}$nsembling), a knowledge fusion framework in which MDLMs iteratively identify reliable decoding trajectories and relay them across models. TIE tracks confidence dynamics over answer-relevant positions to determine which model currently follows a more reliable trajectory and selectively transfers partially denoised sequences across models. As the model on the more promising trajectory often changes across denoising steps, TIE allows different models to contribute complementary strengths at different stages of generation. Strong performance across diverse reasoning tasks, along with our analyses, suggests that TIE offers a practical approach to the underexplored problem of MDLM ensembling.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.