CalibForge: 적대적 솔버 교정을 통한 학습 가능한 최종 작업 확장
CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks
최종 에이전트 훈련에는 단순하게 풀 수 있는 것 이상의, 적절한 수준의 난이도를 가진 실행 가능하고 검증 가능한 작업이 필요합니다. 실행 가능한 검증은 실현 가능성을 확인하지만, 특정 솔버 설정에 대한 작업의 동작 방식을 알려주지는 않습니다. 본 논문에서는 검증된 솔버 동작을 활용하여 적대적 솔버 교정을 통해 후보 작업을 수정하는 자율적인 최종 작업 생성 시스템인 CalibForge를 제시합니다. 다중 솔버 교정은 이기종 솔버 풀 내의 불일치를 목표로 하고, 반면 대비 솔버 교정은 지정된 성공/실패 관계를 목표로 하며, 둘 다 검증된 해결 가능성을 기준으로 하는 솔버 상대적 학습 영역을 구현합니다. CalibForge를 사용하여 5,431개의 교정된 최종 작업을 생성했습니다. 우리의 분석 결과에 따르면, 두 가지 전략 모두 작가 및 검증만 사용하거나 일반적인 단일 솔버 피드백보다 더 효과적인 지도 방식을 제공합니다. 전체 데이터셋으로 학습된 모델은 Terminal-Bench 2.0에서 각각 32.58%와 47.57%의 성능을 달성했습니다. 해당하는 기본 모델 대비 가장 큰 성능 향상은 Terminal-Bench 2.0에서 24.71%, SWE-bench Pro에서 27.68%, Doc2Repo에서 30.04%에 달했습니다. 이러한 결과는 솔버 상대적 학습 가능성이 효과적이고 전이 가능한 에이전트 훈련 데이터 구축을 위한 실용적인 목표임을 뒷받침합니다.
Training terminal agents requires executable and verifiable tasks that are not merely solvable, but appropriately challenging for learning. Executable validation establishes feasibility, yet does not reveal how a task behaves relative to a given solver setting. In this paper, we present CalibForge, an autonomous terminal-task synthesis system that uses verified solver behavior to revise candidate tasks through adversarial solver calibration. Multi-solver calibration targets disagreement within a heterogeneous solver pool, whereas contrastive solver calibration targets a designated strong-pass/weak-fail relation; both operationalize a solver-relative learnable zone anchored in demonstrated solvability. Using CalibForge, we construct 5,431 calibrated terminal tasks. Our ablations show that both strategies yield more effective supervision than authoring and validation alone or ordinary single-solver feedback. Models trained on the full collection achieve 32.58% and 47.57% on Terminal-Bench 2.0. The largest improvements over the corresponding base model reach 24.71 percentage points on Terminal-Bench 2.0, 27.68 points on SWE-bench Pro, and 30.04 points on Doc2Repo. Together, these results support solver-relative learnability as a practical target for constructing effective and transferable agent training data.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.