잔여 오차의 벤치마킹: 매칭된 단기 작업 성능을 넘어 장기 평가가 제공하는 추가적인 정보
Benchmarking the Residual: What Long-Horizon Evaluations Add Beyond Matched Short-Task Performance
장기 벤치마크는 종종 에이전트가 과제가 길어질수록 더 많이 실패한다는 것을 보여줍니다. 이러한 관찰은 실제 적용에 유용하지만, 왜 실패가 발생하는지에 대한 근본적인 설명을 제공하지는 않습니다. 더 많은 단계는 일반적인 오류가 누적될 기회를 늘리고, 더 긴 작업에는 더 어려운 개별 결정이 포함되거나 대화 기록, 도구 출력 및 환경 변화가 축적됨에 따라 어려워질 수 있습니다. 우리는 이러한 마지막 가능성을 '경로 유발 성능 저하'라고 정의합니다. 즉, 초기 실행이 후속 작업을 더욱 어렵게 만듭니다. 모델이 인지하는 텍스트의 누적적인 악영향을 '컨텍스트 부패(context rot)'라고 하는 경우가 많습니다. 본 논문에서는 '장기 실패'를 주장하기 위해서는 벤치마크가 실제 전체 작업 성공률을, 짧은 개별 단계로부터 구축된 기준 예측과 비교해야 한다고 주장합니다. 이러한 기준 예측과 실제 성공률 간의 비율 차이를 '호라이즌 잔여(horizon residual)'라고 부릅니다. 이 비교는 동일한 에이전트 구성을 사용해야 하며, 단계, 체크포인트, 정보 및 예산 선택 방법은 사전에 명시되어야 합니다. 잔여 값은 전체 실행 결과가 선택된 기준과 어떻게 다른지를 보여주며, 왜 이러한 차이가 발생하는지에 대한 추가적인 실험이 필요합니다.
Long-horizon benchmarks often show that agents fail more as tasks become longer. This observation is useful for deployment, but it does not by itself explain why failure occurs. More stages create more opportunities for ordinary errors to compound; longer tasks may also contain harder individual decisions or become harder as conversation history, tool outputs, and environment changes accumulate. We use trajectory-induced degradation to mean this last possibility: earlier execution makes later work harder. When the harmful accumulation is specifically the text visible to the model, it is often called context rot. In this position paper, we argue that to claim a "long-horizon failure", benchmarks must compare actual full-task success against a baseline prediction built from short, individual stages. We call the log-ratio between this prediction and actual success the horizon residual. The comparison must use the same agent configuration and specify in advance how stages, checkpoints, information, and budgets will be chosen. The residual shows that the full rollout differs from the chosen baseline; targeted experiments are still needed to explain why.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.