2607.18960v1 Jul 21, 2026 cs.LG

SFGA: 통계 기반 게이팅 아키텍처 - 신뢰성 있는 SFT 데이터 확보를 위한 심판 단계 증진

SFGA: A Statistics-First Gating Architecture with Adjudicative Escalation for Trustworthy SFT Data Procurement

Arther Tian
Arther Tian
Citations: 7
h-index: 2
Alex Ding
Alex Ding
Citations: 1
h-index: 1
Simon Wu
Simon Wu
Citations: 4
h-index: 1
Aaron Chan
Aaron Chan
Citations: 95
h-index: 4

지도 학습 미세 조정(SFT) 데이터를 확보하는 과정에서 구매자는 하위 작업 훈련 전에 후보 데이터셋이 확보 가치가 있는지 판단해야 합니다. 본 논문에서는 통계 기반 게이팅 아키텍처인 exttt{SFGA}를 제안합니다. exttt{SFGA}는 데이터 확보를 비용을 고려한 라우팅 문제로 간주하며, 다양성, 유용성, 중복성이라는 세 가지 고유한 품질 측면을 활용합니다. 저렴하고 간단한 측정값을 통해 각 축에 대한 추정치를 신뢰 구간과 함께 산출합니다. 게이트는 신뢰 구간이 좁고, 표본 크기가 충분하며, 모든 축의 결과가 일치할 때만 데이터를 수락하고, 그렇지 않은 경우 구매를 옹호하는 심판과 거부를 옹호하는 심판 간의 논쟁을 통해 결정을 내립니다. 독립적인 최종 결정권자가 판결을 내립니다. 통제된 환경에서 12개의 데이터셋(세 가지 축에 대한 $2 imes 3 imes 2$ 격자)과 5개의 시드를 사용하여, 본 게이트는 90%의 정확도와 83%의 F1 점수를 달성했으며, 단위당 비용은 0.017달러입니다. 이는 항상 검증하는 기준(75%)보다 높고, 이상적인 성능(98%)보다는 낮지만, 항상 논쟁으로 이어지는 방식(0.020달러)보다 저렴합니다. 또한, 논쟁 과정에 대한 부정적인 진단 결과를 제시합니다. 거부 옹호 심판의 승률이 80%이고(p ≈ $3 imes 10^{-6}$), 심판 교체 시 52%의 의견 변화율은 기존 LLM 기반 심판이 놓칠 수 있는 편향성을 드러냅니다. 본 연구에서는 주입된 변수를 활용하여 측정 정확성과 라우팅 보정을 위한 통제된 합성 벤치마크를 구성하고, 외부 적용 가능성은 향후 연구 과제로 남겨둡니다.

Original Abstract

Procuring supervised fine-tuning (SFT) data forces a buyer to decide, before any downstream training, whether a candidate corpus is worth acquiring. We present \sys{}, a statistics-first gating architecture that treats procurement as a cost-aware routing problem over three intrinsic quality axes -- diversity, utility, and redundancy. Cheap blind measurements are summarised into per-axis estimates with confidence intervals; a gate accepts a decision only when intervals are tight, sample sizes are adequate, and the axes agree, otherwise it escalates the case to an adjudicative debate between a buy-advocate and a reject-advocate judge, resolved by a presiding verdict. On a controlled benchmark of 12 datasets ($2{\times}3{\times}2$ grid over the three axes) with 5 seeds, the gate reaches 0.90 accuracy and 0.83 $F_1$ at \$0.017 per unit, sitting between an always-verify baseline (0.75) and an oracle upper bound (0.98) while spending less than always-escalate (\$0.020). We further report honest negative diagnostics of the debate path: a con-side win rate of 0.80 ($p\approx3{\times}10^{-6}$) and a 52\% position-flip rate under advocate swapping expose negativity and positional biases that a naive LLM-judge would hide. We frame the injected-knob evaluation explicitly as a controlled synthetic benchmark for measurement fidelity and routing calibration, and delimit external validity as future work.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!