EvoPolicyGym: 상호 작용 환경에서의 자율 정책 진화 평가
EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments
자율 에이전트가 피드백을 통해 실행 가능한 정책을 개선하는 것은 점점 더 중요해지고 있지만, 기존의 평가는 이러한 과정을 종종 최종 점수로 축소하거나 개방형 소프트웨어 엔지니어링 발전과 혼동합니다. 본 논문에서는 제어된 평가 환경인 Autonomous Policy Evolution(자율 정책 진화)을 소개합니다. 이 환경에서 에이전트는 고정된 상호 작용 예산 내에서 반복적으로 실행 가능한 정책 시스템을 수정합니다. 우리는 이를 EvoPolicyGym이라는 벤치마크에 구현하여, 에이전트가 탐색하는 정책을 어떻게 반복적으로 개선하는지 평가합니다. EvoPolicyGym 데이터 세트에서 GPT-5.5는 가장 높은 집계 순위 점수를 달성했으며, 16개 환경 모두에서 상위 두 자릿수 성능을 보였습니다. 리더보드 결과 외에도, EvoPolicyGym은 에이전트가 예산을 어떻게 할당하고 피드백을 매개변수 조정으로 변환하는지 보여주는 세부적인 진단 정보를 제공합니다. 이러한 분석 결과는 강력한 자율 정책 진화가 단순히 특정 작업에서의 승리뿐만 아니라, 작업에 적합한 메커니즘을 발견하고 제한된 피드백 하에서 정책을 개선하는 능력에 달려 있음을 보여줍니다.
Autonomous agents are increasingly expected to improve executable policies through feedback, yet existing evaluations often collapse this process into a final score or confound it with open-ended software-engineering progress. We introduce Autonomous Policy Evolution, a controlled evaluation setting in which a harness-model agent repeatedly edits an executable policy system under a fixed interaction budget. We instantiate this setting in EvoPolicyGym, a benchmark built from compact interactive RL environments that evaluates how agents iteratively improve explored policies. On the EvoPolicyGym suite, GPT-5.5 achieves the strongest aggregate rank score and top-two performance on all 16 environments. Beyond leaderboard results, EvoPolicyGym also provides trajectory-level diagnostics that distinguish how agents allocate budget, convert feedback into parametric tuning. These analyses show that strong autonomous policy evolution depends not only on isolated task wins, but on discovering task-appropriate mechanisms and refining policies under bounded feedback.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.