2608.05519v1 Aug 06, 2026 cs.AI

EcoAgent-Bench: 예산 제약 하의 LLM 에이전트의 경제적 의사 결정 평가

EcoAgent-Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents

Ming Gong
Ming Gong
Citations: 28
h-index: 2
Jie Wu
Jie Wu
Citations: 81
h-index: 5
Feixiang Cheng
Feixiang Cheng
Citations: 9
h-index: 1
Qin Zhao
Qin Zhao
Citations: 0
h-index: 0

에이전트 벤치마크는 일반적으로 작업 완료 여부를 측정하고, 자원 사용량은 부가적인 통계 자료로 취급합니다. 하지만 실제 환경에서는 로컬 검색, 광범위 검색, 복합 연구 도구 사용, 더 강력한 모델 선택 또는 인간 개입 등 다양한 옵션 중에서 선택하는 것이 과제의 일부입니다. 본 논문에서는 각 작업에 가격이 책정된 행동과 명시적인 예산이 설정된 EcoAgent-Bench를 소개합니다. 이 벤치마크는 GAIA, HotpotQA 및 MuSiQue에서 파생된 304개의 실제 기반 작업을 포함하며, 불필요한 에스컬레이션 회피, 로컬 증거가 부족할 때의 에스컬레이션, 모델 계층 선택, 그리고 근거 없는 가정으로 인한 작업 중단 등 네 가지 의사 결정을 평가합니다. 우리는 도구 API 및 워크스페이스 CLI 환경에서 7개의 LLM 에이전트와 4가지 오라클 스크립트 제어 방식을 함께 평가했습니다. 마이크로 평균 정확도는 일방적인 정책을 선호하는 경향을 보입니다. 항상 에스컬레이션을 수행하는 방식은 높은 마이크로 성공률을 달성하지만, 자원 절약을 목표로 하는 작업에서는 실패합니다. 따라서 우리는 정확도와 경제적 일관성(업그레이드 지향 및 절약 지향 그룹에서의 정확도 중 낮은 값)을 함께 보고 이러한 문제점을 드러냅니다. 도구 API 에이전트는 3.9%에서 24.0%의 마이크로 엄격 성공률(최대 7.3%의 경제적 일관성)을 보이며, 종종 정당한 에스컬레이션 전에 작업을 중단하거나 저렴한 작업에 과도한 비용을 지출하는 경향이 있습니다. GPT-5.4의 에스컬레이션 비율은 예산 임계값 변경 시 0%에서 3%로 감소합니다. 이러한 결과는 예산 내에서의 작업 완료와 경제적인 행동 선택이 서로 다른 특성임을 보여줍니다. 우리는 이 두 가지 측면을 연구하는 데 필요한 작업 모음, 변환 파이프라인, 고정된 평가 환경 및 무결성을 보장하는 결과 아티팩트를 공개합니다.

Original Abstract

Agent benchmarks usually measure task completion and treat resource use as an auxiliary statistic. In deployment, however, the choice among a local lookup, broad search, composite research tool, stronger model, or human escalation is part of the task itself. We introduce EcoAgent-Bench, in which every task specifies priced actions and an explicit budget. Its 304 real-derived tasks span five families adapted from GAIA, HotpotQA, and MuSiQue, and test four decisions: avoiding unnecessary escalation, escalating when local evidence is insufficient, selecting a model tier, and stopping on unsupported premises. We evaluate seven LLM agents in tool-API and workspace-CLI settings, together with four oracle scripted controls. Micro-averaged accuracy rewards one-sided policies: always-escalate controls achieve high micro success while failing save-oriented tasks. We therefore also report an economic-consistency score (the worse of accuracy on upgrade-oriented and save-oriented family groups) which exposes this failure. Tool-API agents attain only 3.9-24.0% micro strict success (at most 7.3% economic consistency), often either stopping before warranted escalation or overspending on cheap tasks. A threshold-crossing budget sweep changes GPT-5.4's escalation rate from 0% to only 3%. These results show that completion under a budget and economical action selection are distinct properties. We release the task bundle, transformation pipeline, frozen evaluation environments, and integrity-bound result artifacts needed to study both.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!