OpenThoughts-Agent: 에이전트 모델을 위한 데이터 레시피
OpenThoughts-Agent: Data Recipes for Agentic Models
에이전트 언어 모델은 AI의 활용 범위를 크게 확장하지만, 광범위한 능력을 갖춘 에이전트를 위한 학습 데이터를 어떻게 구축하는지에 대한 공개적인 정보는 아직 부족합니다. SWE-Smith, SERA 및 Nemotron-Terminal과 같은 기존의 공개 노력들은 일반적으로 특정 벤치마크를 대상으로 하므로, 다양한 에이전트 작업에 걸쳐 일반화되는 모델을 어떻게 훈련할 것인가라는 질문에 대한 해답을 제공하지 못했습니다. OpenThoughts-Agent (OT-Agent) 프로젝트는 에이전트 모델 훈련을 위한 완전한 오픈 데이터 관리 파이프라인을 통해 이러한 격차를 해소하고자 합니다. 우리는 파이프라인의 각 단계를 체계적으로 조사하기 위해 100가지 이상의 통제된 실험을 수행했으며, 이를 통해 작업 소스 및 다양성의 중요성에 대한 통찰력을 얻었습니다. 그런 다음, 우리 파이프라인에서 수집한 10만 개의 예시로 구성된 학습 데이터 세트를 구축하고 Qwen3-32B 모델을 이 데이터 세트로 미세 조정했습니다. 그 결과, 7가지 에이전트 벤치마크에서 평균 정확도가 44.8%를 달성했으며, 이는 가장 강력한 기존 공개 데이터 기반 에이전트 모델(Nemotron-Terminal-32B, 40.9%)보다 3.9%p 향상된 수치입니다. 또한, 우리의 학습 데이터는 뛰어난 확장성을 보여주며, 컴퓨팅 자원을 제어하여 비교했을 때 다른 공개 데이터 세트보다 모든 학습 데이터 크기에서 더 우수한 성능을 보였습니다. 우리는 openthoughts.ai 웹사이트를 통해 우리 학습 데이터 세트, 데이터 파이프라인, 실험 데이터 및 모델을 공개적으로 제공함으로써 에이전트 모델 훈련에 대한 향후의 개방적인 연구를 지원하고자 합니다.
Agentic language models dramatically expand the applications of AI yet little is publicly known about how to curate training data for broadly capable agents. Existing open efforts such as SWE-Smith, SERA, and Nemotron-Terminal typically target a single benchmark, leaving open the question of how to train models that generalize across diverse agentic tasks. The OpenThoughts-Agent (OT-Agent) project addresses this gap with a fully open data curation pipeline for training agentic models. We conduct more than 100 controlled ablation experiments to systematically investigate each stage of the pipeline, yielding insights on the importance of task sources and diversity. We then assemble a training set of 100K examples from our pipeline and fine-tune Qwen3-32B on this dataset, which yields an average accuracy of 44.8% across seven agentic benchmarks and a 3.9 percentage point improvement over the strongest existing open data agentic model (Nemotron-Terminal-32B, 40.9%). Moreover, our training data exhibits strong scaling properties, outperforming alternative open datasets at every training set size in compute-controlled comparisons. We publicly release our training sets, data pipeline, experimental data, and models at openthoughts.ai to support future open research on agentic model training.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.