2606.05625v1 Jun 04, 2026 cs.AI

자기 확신 지연 시간: 프롬프트 기반의 암묵적 해킹을 탐지하는 보상 없는 방법

Self-Commitment Latency: A Reward-Free Probe for Prompted Implicit Hacking

Bonan Shen
Bonan Shen
Citations: 8
h-index: 1
Dingyan Shang
Dingyan Shang
Citations: 0
h-index: 0
Tao Ning
Tao Ning
Citations: 6
h-index: 2
You Wang
You Wang
Citations: 52
h-index: 2

언어 모델의 사고 과정이 겉보기에는 양호해 보이더라도, 암묵적인 보상 해킹은 감사하기 어려울 수 있습니다. 최종 답변이 프롬프트 단축키에 의해 결정되었지만, 작성된 추론 과정은 일반적인 문제 해결과 유사하게 보일 수 있기 때문입니다. 검증기 기반의 탐지 방법은 높은 보상을 얻는 초기 단계의 잘린 추론 문맥을 측정하여 이러한 현상을 드러내지만, 작업 특정 보상 신호가 필요합니다. 본 논문에서는 자기 확신 지연 시간이라는 보다 약한 입력 기반의 대안을 제안합니다. 이는 프롬프트된 추론 문맥이 모델 자체의 최종 답변에 얼마나 빨리 확신하는지를 측정합니다. Qwen2.5-3B-Instruct-4bit 모델을 사용하여 통제된 GSM8K 환경에서, 일반적인 프롬프트와 답변 힌트를 포함하는 프롬프트를 비교하여 제안된 탐지 방법을 평가했습니다. 힌트가 포함된 문맥은 정직한 문맥보다 훨씬 빠르고 불확실성이 낮은 수준으로 최종 답변에 확신하는 경향을 보였습니다. 주요 지연 시간 지표인 임계값 0.8에서의 최초 확신 지연 시간은 AUROC 0.878을 달성했으며, 전체 곡선 요약 결과는 확신 범위에서 AUROC 0.926, 평균 미확신 값에서 AUROC 0.904를 달성했습니다. 두 가지 프롬프트 조건 모두 정답을 제공할 때 이 신호는 더 강력하며, 임계값 변화에 따른 안정성을 보입니다. 이러한 결과는 프롬프트 단축키가 사용 가능한 추론 문맥이 보상 모델, 외부 평가자 또는 학습된 분류기 없이도 감지할 수 있는 초기 행동적 확신 특징을 나타낼 수 있음을 보여줍니다.

Original Abstract

Implicit reward hacking is hard to audit when a language model's chain of thought appears benign: a final answer may be anchored by a prompt shortcut while the written reasoning still resembles ordinary problem solving. Verifier-based probes expose such behavior by measuring how early truncated reasoning contexts obtain high reward, but require a task-specific reward signal. This paper proposes a weaker-input alternative, self-commitment latency, which measures how early a prompted reasoning context commits to the model's own final answer. We evaluate the probe in a controlled paired GSM8K setting using Qwen2.5-3B-Instruct-4bit, comparing ordinary prompts with prompts that include an answer hint. Hinted contexts commit substantially earlier and with lower uncertainty than honest contexts. The primary latency metric, first-commitment latency at threshold 0.8, reaches AUROC 0.878; supporting whole-curve summaries reach AUROC 0.926 for commitment range and 0.904 for mean uncommitted mass. The signal is stronger when both prompt conditions answer correctly and remains stable across thresholds. These results show that shortcut-available reasoning contexts can leave an early behavioral commitment signature detectable without a reward model, external judge, or trained classifier.

1 Citations
0 Influential
1 Altmetric
6.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!