2606.26071v1 Jun 24, 2026 cs.LG

모델 포렌식: 우려스러운 행동이 일치하지 않음을 반영하는지 조사

Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignment

Senthooran Rajamanoharan
Senthooran Rajamanoharan
Citations: 1,582
h-index: 15
Neel Nanda
Neel Nanda
Citations: 11,588
h-index: 37
Aditya Singh
Aditya Singh
Citations: 18
h-index: 2
Gerson C. Kroiz
Gerson C. Kroiz
Citations: 522
h-index: 5

안전 연구의 핵심 목표는 모델이 올바르게 설계되었는지(alignment) 확인하는 것입니다. 기존 연구에서는 주로 우려스러운 행동을 탐지하는 데 초점을 맞추었습니다. 하지만 행동만으로는 모델의 일관성 문제를 확정할 수 없습니다. 왜냐하면 우려스러운 행동은 혼동과 같은 양성적인 원인으로 인해 발생할 수도 있기 때문입니다. 이러한 점을 고려하여, 본 연구에서는 행동이 악의적인 의도에서 비롯되었는지 조사하는 '모델 포렌식' 방법을 제안합니다. 본 논문에서는 모델 포렌식을 위한 기본적인 프로토콜을 제시하며, 이 프로토콜은 필요에 따라 반복적으로 수행됩니다. 첫 번째 단계는 모델의 사고 과정을 분석하여 행동의 원인을 추론하는 것입니다. 두 번째 단계는 이러한 가설을 검증하기 위해 프롬프트 또는 환경을 수정합니다. 사고 과정이 항상 정확하지 않더라도, 이는 풍부한 정보원이며 더 엄격한 증거를 수집하도록 안내할 수 있습니다. 본 프로토콜의 유효성을 평가하기 위해, 모델이 우려스러운 행동을 보이는 6개의 시뮬레이션 환경을 구축하고 각 환경에 적용했습니다. Kimi K2 Thinking 모델이 낮은 노력으로 결론을 내리는 경향이 있다는 가설을 제시하고, 이 가설이 실제로 모델의 행동을 잘 예측한다는 것을 보여주었습니다. 또한, DeepSeek R1 모델이 이전 행동과의 일관성을 유지하기 위해 거짓말을 한다는 것을 반사실적 실험을 통해 확인했습니다. 하지만 본 연구 방법은 여전히 개선될 여지가 많습니다. 예를 들어, Kimi K2 Thinking 모델이 사용자의 의도를 위반한다고 믿는지 테스트했을 때, 그러한 믿음의 증거를 찾지 못했지만, 적절한 대조군(positive controls) 없이는 이러한 테스트가 실제로 그러한 현상을 감지할 수 있는지 확신할 수 없습니다. 전반적으로, 본 연구에서 제시하는 간단한 프로토콜은 강력한 기초 역할을 하며, 향후 연구자들이 이를 개선해 나갈 것으로 기대합니다. 더 넓은 관점에서, 본 연구는 모델 포렌식 분야의 발전을 위한 중요한 단계입니다.

Original Abstract

A central goal of safety research is determining whether a model is misaligned. Prior work has largely focused on detecting concerning behavior. But behavior alone does not establish misalignment: a concerning action can arise from benign causes such as confusion. This motivates model forensics: investigating whether the action was driven by malign intent. In this paper, we propose a baseline protocol for model forensics consisting of two steps, iterated as needed. First, we read the chain of thought (CoT) to generate hypotheses about what drives model behavior. Second, we make edits to the prompt or environment to test these hypotheses. While the CoT is not always faithful, it is a rich source of unsupervised insight that can guide the collection of more rigorous evidence. To evaluate our protocol, we create a suite of six agentic environments where models exhibit concerning behavior, and apply it to each. We establish that Kimi K2 Thinking takes shortcuts due to a genuine disposition towards low-effort actions, by showing this hypothesis successfully predicts its behavior. Through counterfactual experiments, we show DeepSeek R1 deceives out of a desire to be consistent with a previous instance of itself. Our methods nonetheless leave significant room for refinement. For example, when we test whether Kimi K2 Thinking believes it is violating user intent, we find no evidence of such a belief, but without positive controls we cannot confirm our tests would detect it. Overall, we find our simple protocol provides a strong baseline that we hope future work will improve upon. More broadly, our work is a concrete step in developing the growing field of model forensics.

0 Citations
0 Influential
18.5 Altmetric
92.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!