2607.26929v1 Jul 29, 2026 cs.CL

동일한 증거, 다른 대상: 언어 모델 상태에서 진단적 증거가 인과 질문에 미치는 영향 분석

Same Evidence, Different Target: Decoding How Diagnostic Evidence Bears on Causal Questions from Language-Model States

Zhuoran Li
Zhuoran Li
Citations: 141
h-index: 5
Weiyi Kong
Weiyi Kong
Citations: 23
h-index: 2

동일한 진단 결과는 특정 인과 주장을 뒷받침하거나 반박할 수 있지만, 대상을 구성하는 모집단, 결과, 추정치, 경로 또는 식별 가정의 차이에 따라 다른 인과 질문에는 적용되지 않을 수 있습니다. 증거와 대상이 함께 변수될 때, 올바른 답변은 언어 표현 방식, 어휘 중복 또는 익숙한 진단 패턴을 반영할 수 있으며, 반드시 증거가 인과 질문에 부합하는 것을 의미하지는 않습니다. 본 연구에서는 동일한 진단적 증거를 그대로 사용하면서 인과 대상을 변경한 쌍으로 구성된 프롬프트를 제시합니다. 각 프롬프트는 증거가 인과 질문에 미치는 영향에 따라 '지지(Favors)', '반박(Challenges)', '미해결(Unresolved)' 또는 '부적절한 대상(Wrong Target)'으로 분류됩니다. 두 프롬프트가 모두 정확하게 분류될 때 한 쌍으로 간주합니다. Qwen2.5-7B-Instruct, Qwen3-8B 및 Llama-3.1-8B-Instruct 모델의 마지막 트랜스포머 블록에서 추출한 최종 토큰의 은닉 상태를 분석하기 위해 별도의 개발 데이터 세트에서 학습된 선형 판독기를 사용했습니다. 9가지 진단 유형을 포괄하는 49쌍으로 구성된 주요 벤치마크에서, 균형 잡힌 정확도는 0.654에서 0.659 사이의 값을 보이며, 18~21쌍이 복구되었습니다. 두 명의 독립적인 인간 평가자는 98개의 프롬프트 중 95개(96.9%)에 대해 동일한 레이블을 부여했습니다. 체크포인트 전체에서 균형 잡힌 정확도와 완전한 쌍 복구는 개발 시나리오 그룹을 유지하는 순열 기반의 결과보다 높습니다. Qwen2.5 모델에서는 전체 프롬프트에 대한 균형 잡힌 정확도가 제한된 입력 방식보다 높으며, 두 방식 간 차이에 대한 부트스트랩 신뢰 구간은 모두 0 이상입니다. 평가 대상 진단 유형에서 가져온 개발 예제를 사용하지 않고 학습된 판독기는 21쌍을 복구했으며, 이는 9가지 유형 중 최소 하나 이상의 유형에 대해 해당됩니다. 은닉 상태 판독기는 균형 잡힌 정확도와 복구된 쌍의 측면에서 답변 옵션 로짓과 텍스트 기반 모델보다 우수한 성능을 보입니다. 이러한 결과는 은닉 상태가 진단적 증거가 인과 대상을 지지, 반박 또는 관련이 없는지를 선형적으로 해독할 수 있는 정보를 포함하고 있음을 보여줍니다.

Original Abstract

The same diagnostic result can support or challenge one causal claim yet fail to address another when the claims concern different populations, outcomes, estimands, pathways, or identifying assumptions. When the evidence and target vary together, a correct answer may reflect favorable or adverse wording, lexical overlap, or a familiar diagnostic pattern rather than matching the evidence to the causal question. We introduce paired prompts that repeat the same diagnostic evidence verbatim while changing the causal target. Each prompt is labeled Favors, Challenges, Unresolved, or Wrong Target according to how the evidence bears on the causal question. A pair is recovered only when both prompts are classified correctly. Using linear readouts trained on a separate development set, we analyze the final-token hidden state from the penultimate transformer block of Qwen2.5-7B-Instruct, Qwen3-8B, and Llama-3.1-8B-Instruct. On the 49-pair primary benchmark spanning nine diagnostic families, balanced accuracy ranges from 0.654 to 0.659 and 18-21 pairs are recovered. Two independent human reviewers assigned the same label to 95 of the 98 prompts (96.9%). Across checkpoints, balanced accuracy and complete-pair recovery exceed permutation nulls that preserve development scenario groups. In Qwen2.5, full-prompt balanced accuracy exceeds both restricted inputs, with paired-bootstrap intervals for both differences above zero. Readouts trained without development examples from the evaluated diagnostic family recover 21 pairs, including at least one in each of the nine families. The hidden-state readout exceeds a linear classifier on answer-option logits and text baselines in balanced accuracy and recovered pairs. These results show that the hidden state contains linearly decodable information about whether diagnostic evidence favors, challenges, or fails to address the causal target.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!