Autodata: 고품질 합성 데이터를 생성하는 지능형 데이터 과학자
Autodata: An agentic data scientist to create high quality synthetic data
본 논문에서는 Autodata를 소개합니다. Autodata는 AI 에이전트가 데이터 과학자와 유사하게 작동하여 고품질의 학습 및 평가 데이터를 구축할 수 있도록 하는 일반적인 방법입니다. 우리는 이러한 데이터 과학자 에이전트를 훈련(메타 최적화)하는 방법을 보여주어, 더욱 강력한 데이터를 생성하도록 학습시킵니다. 전체 프레임워크와 구체적인 실용 구현 방식인 'Agentic Self-Instruct'에 대해 설명합니다. 컴퓨터 과학 연구 과제, 법률 추론 과제 및 수학 객체 기반 추론 과제를 포함한 다양한 실험을 통해, 기존의 합성 데이터셋 생성 방법보다 향상된 결과를 얻었습니다. 더욱이, 데이터 과학자 에이전트 자체를 메타 최적화함으로써 성능을 더욱 크게 향상시킬 수 있습니다. 지능형 데이터 생성이 증가된 추론 연산 능력을 활용하여 더 높은 품질의 모델 훈련을 가능하게 합니다. 전반적으로, 본 연구 방향은 AI 데이터 구축 방식을 혁신할 잠재력이 있다고 믿습니다.
We introduce Autodata, a general method that enables AI agents to act as data scientists who build high quality training and evaluation data. We show how to train (meta-optimize) such a data scientist agent, so that it learns to create even stronger data. We describe the overall formulation, and a specific practical implementation, Agentic Self-Instruct. We conduct experiments on computer science research tasks, legal reasoning tasks and reasoning with mathematical objects, where we obtain improved results compared to classical synthetic dataset creation methods. Further, meta-optimizing the data scientist agent itself delivers an even larger performance uplift. Agentic data creation provides a way to convert increased inference compute into higher quality model training. Overall, we believe this direction has the potential to change the way we build AI data.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.