모델 수준 평가만으로는 배포 관련 정렬(Alignment)을 추론할 수 없다
Deployment-Relevant Alignment Cannot Be Inferred from Model-Level Evaluation Alone
머신러닝 분야에서 정렬 평가는 주로 모델 자체에 대한 평가로 이루어져 왔습니다. 널리 사용되는 벤치마크들은 고정된 입력에 대한 모델의 출력, 예를 들어 진실성, 지시사항 준수 또는 쌍대 비교 선호도 등을 평가하며, 이러한 점수는 종종 실제 배포 환경에서의 정렬을 뒷받침하는 근거로 사용됩니다. 본 논문에서는 배포와 관련된 정렬은 모델 수준 평가만으로는 추론할 수 없다고 주장합니다. 정렬 주장은 증거가 수집되는 수준, 즉 모델 수준, 응답 수준, 상호 작용 수준 또는 배포 수준에 따라 구분되어야 합니다. 본 연구는 이러한 주장을 뒷받침하는 두 가지 연구를 제시합니다. 첫째, 11개의 정렬 벤치마크를 분석하고 16개의 벤치마크로 확장한 결과, Cohen's kappa 값이 0.87인 8가지 차원 기준으로 평가한 결과, 조사된 모든 벤치마크에서 사용자에게 제공되는 검증 지원 기능이 부족하며, 프로세스 제어 기능 또한 거의 부재한 것으로 나타났습니다. tau-bench, CURATe, Rifts, Common Ground 등 소수의 상호 작용 벤치마크는 여전히 내용의 폭이 제한적이며, 벤치마크의 구성 방식이 데이터 소스보다 측정 대상에 더 큰 영향을 미칩니다. 둘째, 3개의 최첨단 모델과 4가지 프레임워크를 사용하여 180개의 텍스트를 기반으로 실시한 익명 모델 간 스트레스 테스트 결과, 동일한 검증 프레임워크가 하나의 모델의 검증 지원 기능을 최고 수준으로 끌어올리는 반면, 다른 모델은 거의 변화가 없는 것으로 나타났습니다. 이는 프레임워크의 효과가 모델에 따라 다르다는 것을 보여주며, 앞서 언급한 격차는 모델 수준에서만으로는 해소될 수 없음을 시사합니다. 우리는 시스템 수준의 평가 체계를 제안합니다. 이는 단일 점수 대신 정렬 프로필을 사용하고, 비교 가능한 상호 작용 평가를 위한 고정된 프레임워크 프로토콜을 사용하며, 평가 증거와 실제 배포 주장의 추론적 간격을 명확하게 밝히는 보고서 템플릿을 사용하는 것을 포함합니다.
Alignment evaluation in machine learning has largely become evaluation of models. Influential benchmarks score model outputs under fixed inputs, such as truthfulness, instruction following, or pairwise preference, and these scores are often used to support claims about deployed alignment. This paper argues that deployment-relevant alignment cannot be inferred from model-level evaluation alone. Alignment claims should instead be indexed to the level at which evidence is collected: model-level, response-level, interaction-level, or deployment-level. Two studies support this position. First, a structured audit of eleven alignment benchmarks, extended to a sixteen-benchmark corpus, dual-coded against an eight-dimension rubric with Cohen's kappa = 0.87, finds that user-facing verification support is absent across every benchmark examined, while process steerability is nearly absent. The few interactional benchmarks identified, including tau-bench, CURATe, Rifts, and Common Ground, remain fragmented in coverage, and benchmark construction rather than data source determines what is measured. Second, a blinded cross-model stress test using 180 transcripts across three frontier models and four scaffolds finds that the same verification scaffold raises one model's verification support to ceiling while leaving another categorically unchanged. This shows that scaffold efficacy is model-dependent and that the gap identified by the audit cannot be closed at the model level alone. We propose a system-level evaluation agenda: alignment profiles instead of single scores, fixed-scaffolding protocols for comparable interactional evaluation, and reporting templates that make the inferential distance between evaluation evidence and deployment claims explicit.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.