RSIBench-Data: 재귀적 자기 개선을 위한 데이터 중심 연구의 성능 측정
RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement
재귀적 자기 개선은 모델 실패에 대한 증거를 더 나은 모델로 전환하는 것을 필요로 합니다. 데이터 중심의 사후 학습 연구는 능력 격차 진단, 훈련 데이터 전략 설계 및 검증, 그리고 체크포인트 피드백을 통한 학습을 포함합니다. LLM 에이전트가 이러한 과정을 자동화할 수 있을까요? 기존 벤치마크는 연구 결정과 최적화, 서비스 제공, 평가 및 시스템 구현을 혼합하여 에이전트의 연구 능력을 가리고 있습니다. 우리는 데이터 중심 연구자로서 고정된 사후 학습 스택을 갖춘 LLM 에이전트를 위한 통제된 벤치마크인 RSIBench-Data를 소개합니다. 에이전트는 고정된 대상 모델에 대해 반복적으로 훈련 데이터 전략을 수정하며, 훈련 및 서비스는 Tinker 기반 서비스를 사용하고, 공식적인 평가는 Harbor 및 E2B 샌드박스를 통해 수행되며, 모든 에이전트에 동일한 예산을 할당합니다. 우리는 소프트웨어 엔지니어링, 터미널 사용, 과학적 질문 응답 및 수학 분야의 여섯 가지 벤치마크에서 네 개의 최첨단 에이전트를 평가했습니다. 에이전트는 데이터 중심 연구의 핵심 능력을 보여주었습니다. 58.33%의 경우, 피드백을 통해 전략을 개선하여 첫 번째 유효한 시도보다 더 나은 결과를 얻었습니다. 그러나 이러한 개선은 일관적이지 않습니다. 최상의 관찰된 점수 이후에도 계속되는 검색 중 78.26%는 최종 시도가 낮은 점수로 끝나며, 나머지는 동일한 최고점을 재현합니다. 따라서 잠재적으로 우수한 후보가 실행 초기에 나타날 수도 있지만, 후속 수정으로 인해 실패할 수도 있습니다. 경로 분석 결과, 더 나은 성능을 보이는 실행에서 다음 네 가지 패턴이 관찰되었습니다. 정확한 가설 설정, 검증 기반의 지도 학습, 행동에 부합하는 데이터 사용, 그리고 강력한 체크포인트 유지입니다. 이러한 결과는 현재 에이전트가 유용한 데이터 중심 발견을 할 수 있지만 아직 피드백을 일관된 개선으로 전환할 수 없다는 것을 시사합니다. RSIBench-Data는 재귀적 자기 개선에 필요한 연구 능력을 측정하고 감사할 수 있는 테스트 환경을 제공합니다. 저희 코드를 다음 주소에서 공개합니다: https://github.com/evolvent-ai/RSIBench-Data.
Recursive self-improvement requires turning evidence of model failures into better models. Data-centric post-training research entails diagnosing capability gaps, designing and validating training-data strategies, and learning from checkpoint feedback. Can LLM agents automate this loop? Existing benchmarks entangle research decisions with optimization, serving, evaluation, and systems implementation, obscuring agents' research capability. We introduce RSIBench-Data, a controlled benchmark of LLM agents as data-centric researchers with a fixed post-training stack. Agents iteratively revise training-data strategies for a fixed target model; training and serving use Tinker-backed services, official evaluation runs through Harbor and E2B sandboxes, and budgets are fixed across agents. We evaluate four frontier agents on six benchmarks across software engineering, terminal use, scientific question answering, and mathematics. Agents demonstrate core data-centric research capabilities: in 58.33\% of settings, they improve upon the first valid attempt by refining strategies from feedback. However, improvement is inconsistent. Among searches continuing after the best observed score, 78.26\% end with a lower-scoring final attempt, while the rest only recover the same peak. A strong candidate may therefore appear early or midway through a run even as later revisions fail. Trajectory analysis identifies four patterns in stronger runs: accurate hypotheses, validation-grounded supervision, behavior-aligned data, and preservation of strong checkpoints. These findings suggest that current agents can make useful data-centric discoveries but cannot yet translate feedback into consistent improvements. RSIBench-Data provides a measurable, auditable testbed for the research capabilities required for recursive self-improvement. We open-source our code at https://github.com/evolvent-ai/RSIBench-Data.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.