2608.14212v1 Aug 14, 2026 cs.AI

APTER: 전문가 기반 평가 기준을 활용한 적응형 추가 학습

APTER: Adaptive Post-Training with Expert-Grounded Rubrics

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Zhe Li
Zhe Li
Citations: 17
h-index: 3
Xu-Yao Zhang
Xu-Yao Zhang
Citations: 28
h-index: 3
Bo Zhang
Bo Zhang
Citations: 13
h-index: 2
Jiansheng Cai
Jiansheng Cai
Citations: 11
h-index: 2
Xukai Wang
Xukai Wang
Citations: 9
h-index: 1
Liangqi Li
Liangqi Li
Citations: 1
h-index: 1
Xiaoyu Shi
Xiaoyu Shi
Citations: 0
h-index: 0

대규모 언어 모델이 전문 분야에 진입함에 따라, 단순히 유창한 답변을 생성하는 것을 넘어, 해당 분야의 제약을 만족하고 중요한 증거를 포함하며 완전한 추론 과정을 제공해야 합니다. 기존의 추가 학습 방법은 종종 전체적인 선호도나 결과 수준의 검증에 의존하는 반면, 최근의 평가 기준 기반 방법은 일반적으로 각 쿼리에 대해 독립적으로 평가 기준을 생성합니다. 전문 분야에서 이러한 제약 없는 평가 기준은 중요한 요구 사항을 누락하거나 샘플마다 다르게 적용되어 지속적인 능력 부족 문제를 진단하고 해결하는 데 어려움을 초래할 수 있습니다. 본 논문에서는 APTER(Adaptive Post-Training with Expert-Grounded Rubrics)라는 프레임워크를 제시합니다. 이는 구조화된 전문 지식을 활용하여 전문 분야의 복잡한 추론 능력을 세밀하게 평가, 최적화하고 진단하는 데 도움을 줍니다. 첫째, 전문가 기반 평가 기준 구축은 해당 분야 전문가들이 정의한 기준 프레임워크에서 시작되며, 각 기준은 안정적인 전문 역량을 나타냅니다. APTER는 각 쿼리에 대해 관련 기준을 선택하고 이를 쿼리 수준의 평가 기준으로 변환하여, 정답 없이도 재사용 가능한 전문가 기준을 실행 가능한 쿼리 수준의 지도 학습 신호로 활용합니다. 둘째, 적응형 추가 학습은 평가 기준 판정을 최적화 신호와 기준 수준 진단 신호 모두로 사용합니다. 낮은 점수를 받은 판정 결과를 기준 ID별로 집계하면 지속적인 부족 부분을 파악하고 강화 학습 과정에서 해당 기준을 중심으로 하는 표적 지도 미세 조정을 수행할 수 있습니다. 수학적 추론과 의료 질문 답변에 대한 실험 결과, APTER는 두 분야 모두에서 일관된 성능 향상을 보였습니다. 세 가지 모델 버전에 걸쳐, APTER는 해당 기본 모델 대비 수학 평균 점수를 최대 15.86%, 의료 평균 점수를 최대 8.04% 향상시켰습니다. 코드 및 평가 기준 데이터셋은 https://github.com/AntDT-APTER/APTER 에서 확인할 수 있습니다.

Original Abstract

As large language models enter professional domains, they must satisfy domain constraints, include critical evidence, and provide complete reasoning rather than merely produce fluent responses. Existing post-training methods often rely on holistic preferences or outcome-level verification, while recent rubric-based methods usually generate rubrics independently for each query. In specialized domains, such unconstrained rubrics may omit critical requirements and vary across samples, hindering the diagnosis and targeted repair of persistent capability deficiencies. We propose APTER (Adaptive Post-Training with Expert-Grounded Rubrics), a framework that integrates structured domain knowledge into fine-grained evaluation, optimization, and diagnosis for specialized complex reasoning. First, expert-grounded rubric construction starts from an expert criteria framework built by domain experts, where each criterion represents a stable professional capability. For each query, APTER selects relevant criteria and instantiates them into query-level rubrics linked to their source criteria, turning reusable expert criteria into executable query-level supervision without reference answers. Second, adaptive post-training uses rubric verdicts as both optimization and criterion-level diagnostic signals. Aggregating low-scoring verdicts by criterion ID reveals persistent deficiencies and triggers targeted supervised fine-tuning updates during reinforcement learning. Experiments on mathematical reasoning and medical question answering show consistent gains across both domains. Across three model generations, APTER improves the mathematics and medical averages over the corresponding base models by up to 15.86 and 8.04 points, respectively. Code and rubric datasets are available at https://github.com/AntDT-APTER/APTER.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!