2607.26820v1 Jul 29, 2026 cs.LG

블랙박스 다중 대화 상호작용에서 경로 수준의 안전 위험 예측

Forecasting Trajectory-Level Safety Risks in Black-Box Multi-Turn Interactions

Dezhang Kong
Dezhang Kong
Citations: 126
h-index: 6
Shi Lin
Shi Lin
Citations: 107
h-index: 5
Peng Qian
Peng Qian
Citations: 1,485
h-index: 11
Dinghao Liu
Dinghao Liu
Citations: 0
h-index: 0
Renjie Sun
Renjie Sun
Citations: 0
h-index: 0
Sifan Wu
Sifan Wu
Citations: 0
h-index: 0
Chenpei Wang
Chenpei Wang
Citations: 3
h-index: 1
Xun Wang
Xun Wang
Citations: 87
h-index: 4

대규모 언어 모델(LLM)이 독립적인 어시스턴트에서 자율 에이전트로 진화함에 따라, LLM의 안전성을 확보하려면 점별 위험 평가를 넘어 위험이 어떻게 발생하고 장기적인 경로를 통해 전개되는지 이해해야 합니다. 다중 대화 상호작용에서는 악의적인 의도가 겉보기에 무해한 단계로 분산되어 점진적으로 재구성될 수 있으며, 이는 궁극적으로 안전 실패로 이어질 수 있습니다. 기존의 안전장치는 주로 나타난 위반 사항을 감지하는 데 중점을 두며, 잠재적인 위험 변화를 예측하고 사전 예방을 가능하게 하는 능력은 부족합니다. 이러한 한계를 해결하기 위해, 우리는 LLM 보호를 단일 단계 위반 탐지에서 경로 수준의 위험 예측으로 발전시키는 안전 위험 예측 프레임워크인 Recast를 제안합니다. Recast는 먼저 단기 대화 진행 상황과 장기적인 역사적 맥락 모두에서 관련 위험 증거를 획득하여 다중 스케일 경로 관점을 활용합니다. 그런 다음, 현재의 위험 상태와 그 시간적 역학을 포착하여 복합적인 위험 변화를 모델링합니다. 마지막으로, 인과적 시간 인코더는 잠재적인 위험 변화 패턴을 학습하고 향후 위험 발생 가능성이 있는 단계를 예측하는 분포를 추정합니다. 7가지 위험 범주에 대한 광범위한 실험 결과, Recast는 안전 실패의 88.3%를 평균 2.41 단계 전에 예측하며, 오탐율은 12.3%로 나타났습니다. 이는 경로 수준 예측이 안전 위반 발생 이전에 잠재적인 위험을 식별하는 데 효과적임을 보여줍니다.

Original Abstract

As large language models (LLMs) evolve from standalone assistants into autonomous agents, ensuring their safety requires shifting beyond pointwise risk assessment to understand how risks emerge and unfold over long-horizon trajectories. In multi-turn interactions, malicious intent can be decomposed across seemingly harmless turns and gradually reconstructed through interaction trajectories, eventually resulting in safety failures. Existing safeguards remain largely reactive, detecting manifested violations while lacking the ability to predict latent risk evolution and enable preemptive prevention. To address this limitation, we propose Recast, a safety risk forecasting framework that advances LLM safeguarding beyond turn-level violation detection to trajectory-level risk prediction. Recast first retrieves risk-relevant evidence from both short-term dialogue progression and long-term historical context via a dual-scale trajectory view. It then models compositional risk evolution by capturing the current risk configuration and its temporal dynamics. Finally, a causal temporal encoder learns latent risk evolution patterns and predicts the distribution of future risk emergence turns. Extensive experiments across 7 risk categories show that Recast predicts 88.3% of future safety failures with an average lead time of 2.41 turns, while maintaining a false alarm rate of 12.3%, showcasing the effectiveness of trajectory-level forecasting in identifying emerging risks before safety violations occur.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!