DecoupleMix: 확장 가능한 VLM 데이터 레시피를 위한 분리된 비율 탐색 및 볼록 할당
DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes
비전 언어 모델(VLM)을 위한 데이터 큐레이션이 활발해짐에 따라, 사전 학습 혼합 데이터를 구성하는 공개적인 방법은 여전히 경험적이며 체계적이지 않습니다. 연구자들은 품질 필터를 통과한 데이터 세트를 단순히 쌓고, 직관에 의존하여 도메인 간 비율을 설정하며, 새로운 데이터를 포함할 수 있는 원칙적이고 설명 가능한 기준이 부족합니다. 또한, 최첨단 레시피는 공개되지 않는 경우가 많습니다. 본 연구에서는 데이터 구성을 체계적인 혼합 최적화 문제로 정의하고, 혼합을 두 가지 직교적인 하위 문제로 분리하여 재현 가능한 엔지니어링 분야로 발전시켰습니다. 즉, 기능 간의 비율(inter-class ratio)과 범주 내에서의 비율(intra-class ratio)입니다. 기능 간 할당에는 단일 변수 반복 검색 방법을 사용하고, 범주 내 구성에는 다차원 데이터 세트 수준의 품질 및 난이도 평가 점수를 사용하여 선택을 제약된 볼록 최적화 문제로 정의했으며, 여기서 다양성 목표를 포함합니다. DecoupleMix 프레임워크는 다음 두 가지 중요한 기능을 제공합니다. 첫째, 수집해야 할 데이터를 안내하고, 둘째, 데이터 세트 검증을 통제되고 설명 가능한 실험으로 만듭니다. 실험 결과, 제안하는 방법은 기존의 경험적인 방법에 비해 일관되게 더 우수한 성능을 보였습니다. 또한, 소규모 프록시에서 발견된 최적 비율은 재조정 없이도 더 큰 규모로 원활하게 적용될 수 있습니다. 800억 개의 추가적인 다중 모드 토큰을 사용하여 사전 학습을 진행한 결과, 저희 VLM 모델은 훨씬 더 큰 다중 모드 자원을 사용하여 훈련된 강력한 오픈 소스 모델과 경쟁력 있는 성능을 보였습니다.
While data curation for Vision Language Models (VLMs) is increasingly active, public practice for constructing pretraining mixtures remains largely heuristic: practitioners stack datasets that pass quality filters, set cross-domain ratios by intuition, and lack a principled, attributable criterion for admitting new data, while frontier recipes remain undisclosed. We formulate data construction as a systematic mixture-optimization problem and turn it into a reproducible engineering discipline by decoupling the mixture into two orthogonal sub-problems: inter-class ratios across capabilities and intra-class ratios within a category. For inter-class allocation, we use a single-variable iterative search; for intra-class composition, we apply a multidimensional, dataset-level assessment scoring Quality and Difficulty, and formulate selection as a constrained convex optimization with a diversity objective. The DecoupleMix framework delivers two critical capabilities: guiding what data to collect next and rendering dataset validation a controlled, attributable experiment. Experiments show our approach consistently surpasses heuristic baselines. Moreover, optimal ratios discovered on small-scale proxies transfer seamlessly to larger scales without retuning. Using 80B additional multimodal continue-pretraining tokens, our VLM is competitive with strong open-source models trained with substantially larger multimodal budgets.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.