2608.02358v1 Aug 03, 2026 cs.CL

ScrambleToolBench: 에이전트가 자신의 지도가 다음 단계를 가리키더라도, 철저한 탐색을 수행하는 방법

ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step

Soujanya Poria
Soujanya Poria
Citations: 36,193
h-index: 80
Zhengyuan Liu
Zhengyuan Liu
Citations: 16
h-index: 2
Navonil Majumder
Navonil Majumder
Citations: 9,789
h-index: 33
Nancy F. Chen
Nancy F. Chen
Citations: 34
h-index: 3
Vernon Toh
Vernon Toh
Citations: 91
h-index: 6

자율 에이전트는 문서 없이도 상호 작용만으로 익숙하지 않은 시스템의 동작을 추론하여 개방형 환경에서 안정적으로 작동할 수 있어야 합니다. 그러나 기존 도구 사용 벤치마크는 정적인 환경에서 의미 있는 도구 스키마를 노출시켜 에이전트가 자율적인 발견 대신 사전 지식에 의존하게 만듭니다. 이러한 한계를 해결하기 위해, 우리는 행동 추론을 평가하도록 설계된 인터랙티브 터미널 벤치마크인 ScrambleToolBench를 소개합니다. 의미 있는 단서를 제거하고 지속적인 작업 교육 과정을 적용함으로써, 이 벤치마크는 에이전트가 시행착오를 통해 숨겨진 도구 동작을 완전히 발견하도록 요구합니다. 또한, 이 벤치마크는 지도 오류, 확률적 액션 실패 및 시간 제약과 같은 동적 과제를 도입하여 환경 변화에 따라 에이전트가 가설을 수정하고 적응할 수 있는지 평가합니다. 최첨단 언어 모델의 평가 결과, 초기 발견 성공은 강력한 적응으로 이어지지 않는다는 것을 보여줍니다. 구조적 변화(예: 지도 오류)에 직면하면 에이전트는 순환 추적과 같은 귀납적 전략을 사용하지 못하고, 대신 기존 믿음에 얽매이거나 포괄적인 탐색에 의존합니다. 테스트 시간 동안의 추론 능력 향상은 이러한 비용이 많이 드는 무차별 검색을 더욱 강화할 뿐이며, 귀납적 복구를 가능하게 하지 않습니다. 지속적인 메모리를 에이전트에 장착하면 누적 오류를 줄일 수 있지만, 에이전트는 여전히 구조적 변화를 효율적으로 추론하는 데 어려움을 겪으며, 이는 현재 에이전트의 추론 능력에 존재하는 격차를 강조합니다.

Original Abstract

To operate robustly in open-world environments, autonomous agents should be able to infer the behavior of unfamiliar systems through interaction alone, even in the absence of documentation. However, existing tool-use benchmarks expose semantic tool schemas in static environments, allowing agents to rely on prior knowledge rather than autonomous discovery. To address this limitation, we introduce ScrambleToolBench, an interactive terminal benchmark designed to isolate behavioral reasoning. By removing semantic cues and enforcing a continuous task curriculum, the benchmark requires agents to uncover hidden tool behaviors entirely through trial-and-error interaction. The benchmark further introduces dynamic challenges, including mapping drift, stochastic action failures, and temporal execution windows, to evaluate whether agents can revise and adapt their hypotheses as the environment changes. Our evaluation of state-of-the-art language models reveals that successful initial discovery does not translate into robust adaptation. When faced with structural changes such as mapping drift, agents fail to use deductive strategies such as cycle tracing, and instead exhibit belief inertia or fall back to exhaustive search. Increasing test-time reasoning only amplifies this expensive brute-force search rather than enabling deductive recovery. While equipping agents with persistent memory reduces compounding errors, they remain unable to efficiently infer structural changes, highlighting a gap in current agent reasoning.

0 Citations
0 Influential
30 Altmetric
150.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!