2608.14312v1 Aug 14, 2026 cs.CL

Envs-FORGE: 에이전트 강화 학습을 위한 프론티어 최적화된 보상 기반 환경 합성

Envs-FORGE: Frontier-Optimized Reward-Grounded Environment Synthesis for Agent RL

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Hao Zhou
Hao Zhou
Citations: 293
h-index: 5
Honghao Liu
Honghao Liu
Citations: 1,498
h-index: 3
Cehao Yang
Cehao Yang
Citations: 136
h-index: 5
Jia Li
Jia Li
Citations: 14
h-index: 3
Jian Guo
Jian Guo
Citations: 2,081
h-index: 13
Xiaojun Wu
Xiaojun Wu
Citations: 43
h-index: 4
Xueyuan Lin
Xueyuan Lin
Citations: 68
h-index: 4

터미널 에이전트를 위한 강화 학습(RL)은 신뢰할 수 있는 보상과 적절한 난이도를 갖춘 실행 가능한 훈련 환경이 필요합니다. Few-shot, Self-Instruct 및 Evol-Instruct과 같은 고정된 레시피는 모든 시드에 동일한 프롬프팅 정책을 적용하며, 이는 현재 정책이 더 어렵거나 쉬운 또는 단순히 다른 작업을 통해 이점을 얻을 수 있는 경우에도 마찬가지입니다. 본 논문에서는 Envs-FORGE를 소개합니다. Envs-FORGE는 검증 보상을 기반으로 각 시드에 대한 환경 합성 액션을 결정하는 프롬프팅 정책입니다. Envs-FORGE는 시드 통과율을 추정하고, 목표 학습 경계를 중심으로 6가지 방향의 예측 액션을 평가하며, 각 시드별로 최적화된 정수 선형 계획법(MILP) 문제를 해결하여 환경 생성을 위한 액션을 선택합니다. 선택된 액션은 명령어, 고정 요소, 오라클 솔루션, 테스트 및 Docker 환경을 동기적으로 재작성하는 데 사용되며, 검증된 번들만 RL 훈련에 사용됩니다. 또한, 이 인덱싱된 MILP 형식은 포트폴리오 계획을 위한 선택적 소프트 스킬 커버리지를 지원합니다. Qwen 3.5 35B 모델에서 Envs-FORGE는 tb-core 데이터셋에서 Base 모델 대비 Pass@1 성능을 9.2%p 향상시켰으며 (40.0%에서 49.2%), tb-2.0 데이터셋에서는 6.4%p 향상시켰습니다 (23.0%에서 29.4%). 또한, Envs-FORGE는 가장 강력한 고정 레시피 모델 대비 각각 2.4%p 및 2.1%p 더 높은 성능을 보였습니다. SWE-bench Verified 데이터셋에서는 77.1%의 성능을 달성하여 Base 모델의 73.4%를 능가했으며, 평가된 4B~35B 모델 전반에 걸쳐 tb-core 데이터셋에서 6.8~9.2%p의 성능 향상을 보였습니다. 모든 합성 방법은 100개의 검증된 환경을 생성하며, 2.27M~2.88M 개의 합성 토큰을 사용하므로, 비교는 동일한 다운스트림 훈련 데이터셋 크기와 운영 규모에서 이루어졌습니다. 소스 코드는 https://github.com/DataArcTech/DataArc-SynData-Toolkit/ 에서 확인할 수 있습니다.

Original Abstract

Reinforcement learning (RL) for terminal agents needs executable training environments with reliable rewards and useful difficulty. Fixed recipes such as few-shot, Self-Instruct, and Evol-Instruct apply the same prompting policy to every seed, even when the current policy would benefit from a harder, easier, or simply different task. We present Envs-FORGE, a prompting policy that converts verifier rewards into per-seed environment-synthesis actions. Envs-FORGE estimates seed pass rates, scores six projection--direction actions around a target learning frontier, and solves a per-seed mixed-integer linear program (MILP) to choose the action that conditions generation. The selected action drives synchronized rewriting of the instruction, fixtures, oracle solution, tests, and Docker environment; only gold-verified bundles enter RL training. The indexed MILP form also supports optional soft skill coverage for portfolio planning. On Qwen 3.5 35B, Envs-FORGE improves Pass@1 over Base by 9.2 percentage points on tb-core (40.0% to 49.2%) and 6.4 points on tb-2.0 (23.0% to 29.4%), exceeding the strongest fixed-recipe baseline by 2.4 and 2.1 points. It reaches 77.1% on SWE-bench Verified versus 73.4% for Base, and improves tb-core by 6.8--9.2 points across the evaluated 4B--35B models. All synthesis methods export 100 verified environments and use 2.27M--2.88M synthesis tokens, placing the comparison at the same downstream training-set size and the same operational scale. The source code is available at https://github.com/DataArcTech/DataArc-SynData-Toolkit/.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!