2605.25893v1 May 25, 2026 cs.AI

D² 모니터: 망설임 인지 라우팅을 통한 확산형 대규모 언어 모델의 동적 안전성 모니터링

$D^2$-Monitor: Dynamic Safety Monitoring for Diffusion LLMs via Hesitation-Aware Routing

Baoyuan Wu
Baoyuan Wu
Citations: 109
h-index: 2
Junchi Yu
Junchi Yu
Citations: 62
h-index: 4
Ao Liu
Ao Liu
Citations: 51
h-index: 3
Yupeng Chen
Yupeng Chen
Citations: 35
h-index: 2
James Oldfield
James Oldfield
Citations: 444
h-index: 9
G. Hong
G. Hong
Citations: 19
h-index: 3
Adel Bibi
Adel Bibi
Citations: 774
h-index: 13
Philip H. S. Torr
Philip H. S. Torr
Citations: 75,314
h-index: 111

오토레그레시브 대규모 언어 모델(AR-LLM)의 대안으로 등장한 확산형 대규모 언어 모델(D-LLM)에 대한 안전성 모니터링은 아직 활발히 연구되지 않았습니다. AR-LLM과 달리 D-LLM은 다단계 노이즈 제거 과정을 통해 텍스트를 생성하며, 이 과정에서 표준적인 단일 단계 모니터링 환경에서는 얻을 수 없는 안전 관련 정보를 포함할 수 있는 중간 은닉 표현이 드러납니다. 항상 켜져 있는(always-on) 모니터링에 적합한 경량 프로브의 활용 가능성에 주목하여, 어떤 수준의 신호가 프로브의 성능 저하를 가장 잘 나타내는지를 분석했습니다. 그 결과, 안전 관련 망설임(safety hesitation)이 가장 유용한 정보라는 것을 확인했습니다. 즉, 중간 은닉 상태가 반복적으로 프로브의 의사 결정 경계 내부에 위치하는 현상은 D-LLM의 생성 과정에서 중요한 지표가 됩니다. 이러한 망설임 단계 수는 D-LLM의 생성 경로(trajectory)를 예측하고, 샘플의 난이도를 나타내는 지표로 활용될 수 있습니다. 이러한 분석을 바탕으로, 본 연구에서는 확산형 대규모 언어 모델을 위한 이중 구조 안전 모니터링 시스템인 D² 모니터를 제안합니다. D² 모니터는 경량 프로브를 항상 켜져 있는 모니터로 사용하여 망설임을 추정하고 기본적인 분류 작업을 수행합니다. 망설임 수준이 특정 임계값을 초과하면, 더 표현력이 뛰어나지만 계산 비용이 높은 프로브가 활성화됩니다. 이러한 동적 라우팅 메커니즘은 테스트 시점에 효율적으로 모니터링 리소스를 할당합니다. 3개의 데이터셋(WildguardMix, ToxicChat, OpenAI-Moderation)과 4개의 D-LLM 모델을 사용하여 평가한 결과, D² 모니터는 최첨단 성능을 달성했으며, 컴팩트한 파라미터 규모(≤ 0.85M 파라미터)를 유지하면서 기존의 8가지 기본 방식에 비해 효과성과 효율성의 균형이 가장 뛰어난 것을 확인했습니다.

Original Abstract

Despite the emergence of diffusion large language models (D-LLMs) as an alternative to autoregressive large language models (AR-LLMs), safety monitoring for D-LLMs remains largely unexplored. Unlike AR-LLMs, D-LLMs generate text through a multi-step denoising process, exposing intermediate hidden representations that may contain safety-relevant information unavailable in standard single-step monitoring setups. Motivated by the suitability of lightweight probes for always-on monitoring, we analyze which trajectory-level signals best indicate when such probes are likely to struggle. We find that the most informative signal is safety hesitation: intermediate hidden states repeatedly falling within a small margin of the probe's decision boundary. The number of such hesitation steps in D-LLM's trajectory predicts probe failure effectively, providing a proxy of sample difficulty. Building on this analysis, we propose $D^2$-Monitor, a bi-level safety monitor for D-LLMs. $D^2$-Monitor adopts a lightweight probe as an always-on monitor to jointly estimate hesitation and perform base classification. When the hesitation level exceeds a threshold, a more expressive but computationally heavier probe is activated. This dynamic routing mechanism allocates monitoring resources efficiently at test time. Evaluated on 3 datasets (WildguardMix, ToxicChat, OpenAI-Moderation) across 4 D-LLMs, $D^2$-Monitor achieves state-of-the-art performance with a compact parameter footprint ($\leq$ 0.85M parameters), and exhibits the best trade-off between effectiveness and efficiency relative to 8 baselines.

1 Citations
0 Influential
30 Altmetric
151.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!