기술 활용인가, 기술 연극인가? 기술 강화 언어 에이전트의 추론 과정 평가
Skill Use or Skill Theater? Evaluating the Reasoning Backroom in Skill-Augmented Language Agents
재사용 가능한 기술은 작업 절차를 통해 언어 에이전트를 확장하는 표준 인터페이스로 자리 잡고 있습니다. 그러나 평가자는 일반적으로 가시적인 추론 또는 에이전트 자신의 설명으로부터 기술 활용을 추론합니다. 이러한 신호는 에이전트가 보이는 사용 방식을 나타낼 뿐, 실제로 기술이 의사 결정에 영향을 미쳤는지 여부를 보여주지는 않습니다. 본 연구에서는 기술 강화 에이전트가 '추론 과정의 뒷면(Reasoning Backroom)' 현상을 보이는지 조사합니다. 즉, 명시된 기술 활용과 실제 개입을 통해 측정된 영향 간에 체계적인 격차가 존재하는지를 확인하고자 합니다. 우리는 BACKTRACE라는 평가 프레임워크를 소개하며, 이 프레임워크는 각 기술 기반 답변에 대해 일치하는 '기술 없음(no-skill)' 대조군을 사용하고, 기술의 의미, 표현, 동일성, 내용 및 할당에 개입하며, 답변이 완료된 후에만 설명 정보를 수집합니다. 우리는 이 프레임워크를 BACKROOMBench라는 검증된 테스트베드로 구현했습니다. 이 테스트베드는 통제된 논리 및 경쟁 수학 영역을 포괄하며, 다양한 기술 조건, 단일 에이전트 및 다중 에이전트 환경, 그리고 다양한 모델 아키텍처를 포함합니다. 우리의 평가는 광범위한 출처 불분명 문제를 드러냅니다. 모델과 도메인에 관계없이, 명시된 기술 활용은 종종 안정적으로 유지되는 반면, 인과적 의존성과 유틸리티는 변동하며, 이는 '침묵하는 학습'과 '겉으로 보이는 사용'을 모두 야기합니다. 행동 효과는 표시된 기술의 동일성보다 절차적 내용에 더 일관되게 나타나며, 명시된 설명은 사용 가능한 정보의 양에 크게 영향을 받습니다. 직접적인 기술 활용 주장, 텍스트 언급, 추적 유사성 및 LLM 평가기를 기반으로 한 관찰 기반 탐지기는 실제로 기술에 의존하는 결정을 식별하지 못합니다. 다중 에이전트 시스템에서는 기술의 영향력이 정보 출처가 사라진 후에도 통신을 통해 유지될 수 있으며, '기술 없음' 팀은 제공되지 않은 기술과 출처를 여전히 언급하기도 합니다. 이러한 결과는 추론 과정의 뒷면 현상이 일반적인 AI 출처 불분명 문제이며, 이를 해결하기 위해서는 개입이 필요하다는 것을 보여줍니다.
Reusable skills are becoming a standard interface for extending language agents with task procedures. Yet evaluators usually infer skill use from visible reasoning or the agent's own attribution. These signals show what the agent appears to use, not whether the skill changed its decision. We ask whether skill-augmented agents exhibit a \textbf{Reasoning Backroom}, a systematic gap between stated skill use and intervention-measured influence. We introduce BACKTRACE, an evaluation framework that pairs each skill-conditioned answer with a matched no-skill counterfactual, intervenes on skill meaning, wording, identity, content, and assignment, and elicits attribution only after the answer is committed. We instantiate the framework as BACKROOMBench, a verified testbed spanning controlled logic and competition mathematics, multiple skill conditions, single-agent and multi-agent settings, and diverse model families. Our evaluation reveals a pervasive provenance failure. Across models and domains, stated skill use often remains stable while causal reliance and signed utility vary, producing both silent uptake and performative use. Behavioral effects follow procedural content more reliably than displayed skill identity, whereas stated attributions respond strongly to artifact availability. Observational detectors based on direct skill-use claims, text mentions, trace similarity, and an LLM judge do not identify which decisions actually depend on the skill. In multi-agent systems, skill influence can survive communication even after its source is lost, while no-skill teams still name skills and sources that were never supplied. These findings establish the Reasoning Backroom as a general AI provenance problem whose audit requires intervention.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.