HSS-Synth: LLM을 위한 인문사회 데이터 합성
HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs
고품질의 다양한 데이터는 대규모 언어 모델(LLM)에 필수적이지만, 이러한 데이터는 여전히 부족하고 비용이 많이 듭니다. 데이터 합성은 실행 가능한 대안이며, 특정 작업에서는 성공적인 결과를 보여주지만, 인문사회 분야(HSS)는 간과되어 왔으며, 그 개방적인 특성으로 인해 합성 자체가 어렵습니다. 기존의 기능 중심적이고 단편적인 시도들을 넘어, 우리는 주제 중심적인 패러다임을 채택하고, 14개의 주요 분야를 포괄하는 최초의 HSS 도메인 시스템을 정의하며, 최초의 HSS 데이터 합성 파이프라인인 HSS-Synth를 소개합니다. HSS-Synth는 다음과 같은 구성 요소로 이루어집니다: (1) 다단계 필터링 및 텍스트 정제를 거쳐 웹 코퍼스에서 시드 문서를 구축하고, 심사관에 의해 평가합니다; (2) '요구사항 + 페르소나'를 명시하여 시드 문서를 다양한 방식으로 역번역하고, 엄격한 질의응답 일치성 검사를 수행하여 충실하면서도 다채로운 지침을 생성합니다; (3) LLM 응답 제한을 극복하기 위해, 답변 생성 과정에서 시드 문서를 활용하는 '선생-강요 방식'의 질문 답변 방식을 적용하여 의미를 고정하고, 환각 현상을 줄이며, 어조와 진실성을 유지합니다. HSS-Synth는 237,000개의 고품질, 다양한 지침 조정 샘플을 생성하며, 이는 16개의 벤치마크에서 14개의 주요 모델보다 뛰어난 성능을 보입니다. 미세 조정된 Qwen3-8B-Base 모델은 새로운 최고 수준의 성능을 달성했으며, 공식 Qwen3-8B에 근접하는 동시에 인간 선호도 및 지식 능력을 향상시켰으며, 성능 저하 없이 이러한 개선이 이루어졌습니다. 광범위한 실험 결과는 HSS-Synth의 견고성과 전송 가능성을 입증합니다. 저희 코드는 https://github.com/pengr/HSS-Synth 에서 공개적으로 이용할 수 있습니다.
High-quality, diverse data are vital for large language models (LLMs) but remain scarce and costly. Data synthesis is a viable alternative and succeeds on closed tasks, yet the humanities and social sciences (HSS) are overlooked, and their open-ended nature makes synthesis challenging. Moving beyond prior capability-centric, fragmented attempts, we adopt a subject-centric paradigm, define the first HSS domain system covering 14 mainstream fields, and introduce HSS-Synth, the first data synthesis pipeline for HSS. HSS-Synth comprises: (1) constructing seed documents from web corpora via multi-step filtering and text refinement evaluated by a judge; (2) specifying "requirements + persona" to backtranslate seed documents into diverse yet faithful instructions with a strict Q&A alignment check; and (3) breaking LLM response limits via teacher-forced Answering that feeds seed documents during response generation to anchor semantics, reduce hallucinations, and preserve tone and integrity. HSS-Synth yields 237k high-quality, diverse instruction-tuning samples that outperform 14 leading baselines on 16 benchmarks. The fine-tuned Qwen3-8B-Base sets a new SOTA and approaches the official Qwen3-8B, improving both human preference and knowledge capabilities without performance seesaws. Extensive experiments demonstrate HSS-Synth's robustness and transferability. Our code is publicly available at https://github.com/pengr/HSS-Synth.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.