2607.02440v1 Jul 02, 2026 cs.AI

EvoPolicyGym: 상호 작용 환경에서의 자율 정책 진화 평가

EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments

Jiacheng Chen
Jiacheng Chen
Citations: 455
h-index: 4
Yulun Wu
Yulun Wu
Citations: 67
h-index: 4
Yafu Li
Yafu Li
Citations: 370
h-index: 8
Yu Cheng
Yu Cheng
Citations: 122
h-index: 5
Qingyu Yin
Qingyu Yin
Citations: 150
h-index: 6
Zhennan Shen
Zhennan Shen
Citations: 24
h-index: 2
Zhilin Wang
Zhilin Wang
Citations: 229
h-index: 5
Tong Zhu
Tong Zhu
Citations: 14
h-index: 2
Guanjie Chen
Guanjie Chen
Citations: 3,416
h-index: 6
Runzhe Zhan
Runzhe Zhan
University of Macau
Citations: 592
h-index: 11
Jusen Du
Jusen Du
Citations: 76
h-index: 5
Tianle Li
Tianle Li
Citations: 10
h-index: 2
Derek F. Wong
Derek F. Wong
Citations: 33
h-index: 2
Yang Yang
Yang Yang
Citations: 101
h-index: 4
Hanxiao Song
Hanxiao Song
Citations: 0
h-index: 0
Yanshu Li
Yanshu Li
University of Texas at Austin
Citations: 467
h-index: 10

자율 에이전트가 피드백을 통해 실행 가능한 정책을 개선하는 것은 점점 더 중요해지고 있지만, 기존의 평가는 이러한 과정을 종종 최종 점수로 축소하거나 개방형 소프트웨어 엔지니어링 발전과 혼동합니다. 본 논문에서는 제어된 평가 환경인 Autonomous Policy Evolution(자율 정책 진화)을 소개합니다. 이 환경에서 에이전트는 고정된 상호 작용 예산 내에서 반복적으로 실행 가능한 정책 시스템을 수정합니다. 우리는 이를 EvoPolicyGym이라는 벤치마크에 구현하여, 에이전트가 탐색하는 정책을 어떻게 반복적으로 개선하는지 평가합니다. EvoPolicyGym 데이터 세트에서 GPT-5.5는 가장 높은 집계 순위 점수를 달성했으며, 16개 환경 모두에서 상위 두 자릿수 성능을 보였습니다. 리더보드 결과 외에도, EvoPolicyGym은 에이전트가 예산을 어떻게 할당하고 피드백을 매개변수 조정으로 변환하는지 보여주는 세부적인 진단 정보를 제공합니다. 이러한 분석 결과는 강력한 자율 정책 진화가 단순히 특정 작업에서의 승리뿐만 아니라, 작업에 적합한 메커니즘을 발견하고 제한된 피드백 하에서 정책을 개선하는 능력에 달려 있음을 보여줍니다.

Original Abstract

Autonomous agents are increasingly expected to improve executable policies through feedback, yet existing evaluations often collapse this process into a final score or confound it with open-ended software-engineering progress. We introduce Autonomous Policy Evolution, a controlled evaluation setting in which a harness-model agent repeatedly edits an executable policy system under a fixed interaction budget. We instantiate this setting in EvoPolicyGym, a benchmark built from compact interactive RL environments that evaluates how agents iteratively improve explored policies. On the EvoPolicyGym suite, GPT-5.5 achieves the strongest aggregate rank score and top-two performance on all 16 environments. Beyond leaderboard results, EvoPolicyGym also provides trajectory-level diagnostics that distinguish how agents allocate budget, convert feedback into parametric tuning. These analyses show that strong autonomous policy evolution depends not only on isolated task wins, but on discovering task-appropriate mechanisms and refining policies under bounded feedback.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!