D² 모니터: 망설임 인지 라우팅을 통한 확산형 대규모 언어 모델의 동적 안전성 모니터링
$D^2$-Monitor: Dynamic Safety Monitoring for Diffusion LLMs via Hesitation-Aware Routing
오토레그레시브 대규모 언어 모델(AR-LLM)의 대안으로 등장한 확산형 대규모 언어 모델(D-LLM)에 대한 안전성 모니터링은 아직 활발히 연구되지 않았습니다. AR-LLM과 달리 D-LLM은 다단계 노이즈 제거 과정을 통해 텍스트를 생성하며, 이 과정에서 표준적인 단일 단계 모니터링 환경에서는 얻을 수 없는 안전 관련 정보를 포함할 수 있는 중간 은닉 표현이 드러납니다. 항상 켜져 있는(always-on) 모니터링에 적합한 경량 프로브의 활용 가능성에 주목하여, 어떤 수준의 신호가 프로브의 성능 저하를 가장 잘 나타내는지를 분석했습니다. 그 결과, 안전 관련 망설임(safety hesitation)이 가장 유용한 정보라는 것을 확인했습니다. 즉, 중간 은닉 상태가 반복적으로 프로브의 의사 결정 경계 내부에 위치하는 현상은 D-LLM의 생성 과정에서 중요한 지표가 됩니다. 이러한 망설임 단계 수는 D-LLM의 생성 경로(trajectory)를 예측하고, 샘플의 난이도를 나타내는 지표로 활용될 수 있습니다. 이러한 분석을 바탕으로, 본 연구에서는 확산형 대규모 언어 모델을 위한 이중 구조 안전 모니터링 시스템인 D² 모니터를 제안합니다. D² 모니터는 경량 프로브를 항상 켜져 있는 모니터로 사용하여 망설임을 추정하고 기본적인 분류 작업을 수행합니다. 망설임 수준이 특정 임계값을 초과하면, 더 표현력이 뛰어나지만 계산 비용이 높은 프로브가 활성화됩니다. 이러한 동적 라우팅 메커니즘은 테스트 시점에 효율적으로 모니터링 리소스를 할당합니다. 3개의 데이터셋(WildguardMix, ToxicChat, OpenAI-Moderation)과 4개의 D-LLM 모델을 사용하여 평가한 결과, D² 모니터는 최첨단 성능을 달성했으며, 컴팩트한 파라미터 규모(≤ 0.85M 파라미터)를 유지하면서 기존의 8가지 기본 방식에 비해 효과성과 효율성의 균형이 가장 뛰어난 것을 확인했습니다.
Despite the emergence of diffusion large language models (D-LLMs) as an alternative to autoregressive large language models (AR-LLMs), safety monitoring for D-LLMs remains largely unexplored. Unlike AR-LLMs, D-LLMs generate text through a multi-step denoising process, exposing intermediate hidden representations that may contain safety-relevant information unavailable in standard single-step monitoring setups. Motivated by the suitability of lightweight probes for always-on monitoring, we analyze which trajectory-level signals best indicate when such probes are likely to struggle. We find that the most informative signal is safety hesitation: intermediate hidden states repeatedly falling within a small margin of the probe's decision boundary. The number of such hesitation steps in D-LLM's trajectory predicts probe failure effectively, providing a proxy of sample difficulty. Building on this analysis, we propose $D^2$-Monitor, a bi-level safety monitor for D-LLMs. $D^2$-Monitor adopts a lightweight probe as an always-on monitor to jointly estimate hesitation and perform base classification. When the hesitation level exceeds a threshold, a more expressive but computationally heavier probe is activated. This dynamic routing mechanism allocates monitoring resources efficiently at test time. Evaluated on 3 datasets (WildguardMix, ToxicChat, OpenAI-Moderation) across 4 D-LLMs, $D^2$-Monitor achieves state-of-the-art performance with a compact parameter footprint ($\leq$ 0.85M parameters), and exhibits the best trade-off between effectiveness and efficiency relative to 8 baselines.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.