2606.24855v1 Jun 23, 2026 cs.AI

OpenThoughts-Agent: 에이전트 모델을 위한 데이터 레시피

OpenThoughts-Agent: Data Recipes for Agentic Models

S. Dillmann
S. Dillmann
Citations: 624
h-index: 6
Xiangyi Li
Xiangyi Li
Citations: 122
h-index: 2
Hanwen Xing
Hanwen Xing
Citations: 127
h-index: 3
Atula Tejaswi
Atula Tejaswi
Citations: 83
h-index: 2
E. K. Buchanan
E. K. Buchanan
Citations: 263
h-index: 4
Marianna Nezhurina
Marianna Nezhurina
Citations: 4,035
h-index: 8
J. Jitsev
J. Jitsev
Citations: 11,357
h-index: 19
Robert Zhang
Robert Zhang
Citations: 392
h-index: 6
L. Chen
L. Chen
Citations: 649
h-index: 3
Anurag Kashyap
Anurag Kashyap
Citations: 186
h-index: 2
E. Guha
E. Guha
Citations: 896
h-index: 10
Negin Raoof
Negin Raoof
Citations: 537
h-index: 4
Ryan Marten
Ryan Marten
Citations: 354
h-index: 2
A. Dimakis
A. Dimakis
Citations: 1,179
h-index: 10
Ben Feuer
Ben Feuer
Citations: 1,206
h-index: 12
Xiaokun Chen
Xiaokun Chen
Citations: 244
h-index: 5
Wanjia Zhao
Wanjia Zhao
Citations: 250
h-index: 3
Ludwig Schmidt
Ludwig Schmidt
Citations: 4
h-index: 1
Zhiwei Xu
Zhiwei Xu
Citations: 1
h-index: 1
Xunyi Jiang
Xunyi Jiang
Citations: 10
h-index: 2
Ashima Suvarna
Ashima Suvarna
UCLA
Citations: 345
h-index: 7
Saadia Gabriel
Saadia Gabriel
Citations: 3,894
h-index: 13
Ling Shi
Ling Shi
Citations: 119
h-index: 5
Nishad Singhi
Nishad Singhi
Citations: 340
h-index: 5
Sujay Sanghavi
Sujay Sanghavi
Citations: 12
h-index: 1
Artem Gazizov
Artem Gazizov
Citations: 15
h-index: 2
Charlie Ruan
Charlie Ruan
Citations: 46
h-index: 3
Tyler Griggs
Tyler Griggs
Citations: 492
h-index: 7
A. Shaw
A. Shaw
Citations: 158
h-index: 3
Hritik Bansal
Hritik Bansal
Citations: 4,152
h-index: 25
Reinhard Heckel
Reinhard Heckel
Citations: 181
h-index: 2
Chinmay Hegde
Chinmay Hegde
Citations: 3
h-index: 1
Sankalp Jajee
Sankalp Jajee
Citations: 1
h-index: 1
Daanish Khazi
Daanish Khazi
Citations: 0
h-index: 0
Emmanouil Koukoumidis
Emmanouil Koukoumidis
Citations: 834
h-index: 8
Han Liu
Han Liu
Citations: 186
h-index: 7
Shlok Natarajan
Shlok Natarajan
Citations: 84
h-index: 3
Harsh Raj
Harsh Raj
Citations: 103
h-index: 4
Nicholas Roberts
Nicholas Roberts
Citations: 24
h-index: 3
Ethan Shen
Ethan Shen
University of Washington
Citations: 79
h-index: 3
Michael Siu
Michael Siu
Citations: 25
h-index: 2
Patrick Yubeaton
Patrick Yubeaton
Citations: 36
h-index: 3
Boxuan Li
Boxuan Li
Citations: 990
h-index: 4
Yein Park
Yein Park
Korea University
Citations: 99
h-index: 4
Minh Pham
Minh Pham
Citations: 186
h-index: 3
Ke Sun
Ke Sun
Citations: 3
h-index: 1
Yixin Wang
Yixin Wang
Citations: 0
h-index: 0
E. Zhang
E. Zhang
Citations: 22
h-index: 1
Siyan Zhao
Siyan Zhao
Citations: 584
h-index: 9
Richard Zhuang
Richard Zhuang
Citations: 230
h-index: 6

에이전트 언어 모델은 AI의 활용 범위를 크게 확장하지만, 광범위한 능력을 갖춘 에이전트를 위한 학습 데이터를 어떻게 구축하는지에 대한 공개적인 정보는 아직 부족합니다. SWE-Smith, SERA 및 Nemotron-Terminal과 같은 기존의 공개 노력들은 일반적으로 특정 벤치마크를 대상으로 하므로, 다양한 에이전트 작업에 걸쳐 일반화되는 모델을 어떻게 훈련할 것인가라는 질문에 대한 해답을 제공하지 못했습니다. OpenThoughts-Agent (OT-Agent) 프로젝트는 에이전트 모델 훈련을 위한 완전한 오픈 데이터 관리 파이프라인을 통해 이러한 격차를 해소하고자 합니다. 우리는 파이프라인의 각 단계를 체계적으로 조사하기 위해 100가지 이상의 통제된 실험을 수행했으며, 이를 통해 작업 소스 및 다양성의 중요성에 대한 통찰력을 얻었습니다. 그런 다음, 우리 파이프라인에서 수집한 10만 개의 예시로 구성된 학습 데이터 세트를 구축하고 Qwen3-32B 모델을 이 데이터 세트로 미세 조정했습니다. 그 결과, 7가지 에이전트 벤치마크에서 평균 정확도가 44.8%를 달성했으며, 이는 가장 강력한 기존 공개 데이터 기반 에이전트 모델(Nemotron-Terminal-32B, 40.9%)보다 3.9%p 향상된 수치입니다. 또한, 우리의 학습 데이터는 뛰어난 확장성을 보여주며, 컴퓨팅 자원을 제어하여 비교했을 때 다른 공개 데이터 세트보다 모든 학습 데이터 크기에서 더 우수한 성능을 보였습니다. 우리는 openthoughts.ai 웹사이트를 통해 우리 학습 데이터 세트, 데이터 파이프라인, 실험 데이터 및 모델을 공개적으로 제공함으로써 에이전트 모델 훈련에 대한 향후의 개방적인 연구를 지원하고자 합니다.

Original Abstract

Agentic language models dramatically expand the applications of AI yet little is publicly known about how to curate training data for broadly capable agents. Existing open efforts such as SWE-Smith, SERA, and Nemotron-Terminal typically target a single benchmark, leaving open the question of how to train models that generalize across diverse agentic tasks. The OpenThoughts-Agent (OT-Agent) project addresses this gap with a fully open data curation pipeline for training agentic models. We conduct more than 100 controlled ablation experiments to systematically investigate each stage of the pipeline, yielding insights on the importance of task sources and diversity. We then assemble a training set of 100K examples from our pipeline and fine-tune Qwen3-32B on this dataset, which yields an average accuracy of 44.8% across seven agentic benchmarks and a 3.9 percentage point improvement over the strongest existing open data agentic model (Nemotron-Terminal-32B, 40.9%). Moreover, our training data exhibits strong scaling properties, outperforming alternative open datasets at every training set size in compute-controlled comparisons. We publicly release our training sets, data pipeline, experimental data, and models at openthoughts.ai to support future open research on agentic model training.

1 Citations
0 Influential
12.5 Altmetric
63.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!