2606.16276v1 Jun 15, 2026 cs.AI

SpecAlign: 합성 데이터를 활용한 대규모 언어 모델의 효율적인 명세 기반 정렬 방법

SpecAlign: Efficient Specification-Grounded Alignment of Large Language Models via Synthetic Data

Xiangliang Zhang
Xiangliang Zhang
Citations: 997
h-index: 16
Zhengqing Yuan
Zhengqing Yuan
Citations: 976
h-index: 10
Wenjie Wang
Wenjie Wang
Citations: 69
h-index: 4
Yue Zhao
Yue Zhao
Citations: 7
h-index: 2
Yue Huang
Yue Huang
Citations: 662
h-index: 10
Shiyi Du
Shiyi Du
Citations: 47
h-index: 2
Han Bao
Han Bao
Citations: 93
h-index: 4
Yuchen Ma
Yuchen Ma
Citations: 70
h-index: 4
Yanfang Ye
Yanfang Ye
Citations: 193
h-index: 7

대규모 언어 모델(LLM)이 실제 응용 분야에 점점 더 많이 사용됨에 따라, 정렬은 단일하고 보편적인 안전성 또는 유용성의 개념에 의해 결정되는 것이 아니라, 제공업체 또는 애플리케이션별 모델 명세에 따라 달라집니다. 이러한 명세는 일반적으로 길고 구조화되어 있으며 자주 업데이트되지만, 기존의 정렬 파이프라인은 이러한 명세를 훈련 신호로 활용할 수 있는 체계적인 메커니즘을 갖추지 못하고 있습니다. 본 논문에서는 제공업체가 작성한 모델 명세를 추상적인 원칙이나 정적 벤치마크가 아닌 주요 정렬 대상으로 간주하는 새로운 정렬 패러다임인 명세 기반 정렬을 제안합니다. 이 패러다임을 구현하기 위해, 우리는 명세 문서에서 직접 정렬 데이터를 생성하는 프레임워크인 SpecAlign을 소개합니다. SpecAlign은 구조화된 규칙 주석, 제어 가능한 명세 인스턴스화 및 다중 에이전트 적대적 데이터 합성을 결합하여 규정 준수 행동과 의미 있는 명세 위반 모두를 포착하는 세밀하고 경계 인식적인 선호도 쌍을 생성합니다. 여러 모델 명세와 기반 모델에 대한 실험 결과, SpecAlign으로 훈련하면 규칙 준수율이 꾸준히 향상되는 동시에 일반적인 능력은 유지되고 지나치게 보수적인 행동은 피할 수 있음을 보여줍니다. 이러한 결과는 명시적인 모델 명세를 기반으로 한 정렬이 LLM의 동작을 변화하는 정책 요구 사항에 맞춰 신속하고 정확하며 확장 가능하게 조정할 수 있도록 한다는 것을 시사합니다.

Original Abstract

As large language models (LLMs) are increasingly deployed in real-world applications, alignment is no longer governed by a single universal notion of safety or helpfulness, but instead by provider- or application-specific model specifications. These specifications are typically long, structured, and frequently updated, yet existing alignment pipelines lack a systematic mechanism to operationalize them as training signals. In this paper, we propose specification-grounded alignment, a new alignment paradigm that treats provider-authored model specifications as the primary alignment target rather than abstract principles or static benchmarks. To instantiate this paradigm, we introduce SpecAlign, a framework that synthesizes alignment data directly from specification documents. SpecAlign combines structured rule annotation, controllable specification instantiation, and multi-agent adversarial data synthesis to generate fine-grained, boundary-aware preference pairs that capture both compliant behaviors and meaningful specification violations. Experiments across multiple model specifications and backbone models demonstrate that training with SpecAlign consistently improves rule compliance while preserving general capabilities and avoiding over-conservative behavior. These results suggest that grounding alignment in explicit model specifications enables rapid, precise, and scalable adaptation of LLM behavior to evolving policy requirements.

1 Citations
0 Influential
8 Altmetric
41.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!