헌법 기반 선호도 재구성 분야의 미해결 과제
Open Problems in Constitutional Preference Reconstruction
쌍대 비교(pairwise) 선호도 데이터는 언어 모델 학습 및 평가에 널리 사용되지만, 각 데이터 포인트는 선택 사항을 기록할 뿐, 그 이면에 있는 근거를 나타내지 않습니다. Inverse Constitutional AI (ICAI)와 같은 방법은 데이터 세트를 짧은 자연어 원칙으로 압축하여 해석 가능성을 높이려고 시도합니다. 우리는 이러한 접근 방식이 충분한 정보를 담고 있지 않다고 주장합니다. 단순히 나열된 원칙들은 실행 가능한 의사 결정 규칙이 될 수 없기 때문입니다. 이 연구에서는 쌍대 비교 설정을 활용하여 헌법 기반 방법론에서 발생하는 세 가지 주요 미해결 과제를 실증적으로 분석했습니다. 첫째, 원칙의 품질을 측정하기 어렵습니다. 커버리지와 정확성은 전체적인 재구성 과정을 평가하는 데 유용하지만 불완전한 지표입니다. 둘째, 원칙 간의 조합이 모호합니다. 원칙을 고정했을 때, 서로 다른 실행자(LLM 판단자와 다수결)가 일치하는 비율은 73%에 불과합니다. 셋째, LLM 모델마다 헌법 내용이 달라집니다. 모델 간 일치율은 73%인 반면, 동일 모델 내에서 일치율은 81%입니다. PRISM, AlpacaEval, Chatbot Arena를 분석한 결과, 원칙 개선(ICAI+)이 이러한 문제들을 해결하는 첫걸음이 될 수 있음을 확인했습니다. 원칙 개선을 통해 실행자 간 일치율이 78%로 상승했으며, 투명한 실행자는 LLM 판단자와 유사한 정확도(66% vs. 67%)를 보였습니다. 본 연구 결과는 헌법을 단순히 원칙의 집합으로 평가하는 것이 아니라, '헌법-실행자 시스템'으로서 평가해야 함을 시사하며, 이는 LLM을 판단자로 활용하는 데 중요한 의미를 가집니다.
Pairwise preference data is widely used for training and evaluating language models (e.g., RLHF), but each datapoint records a \emph{choice}, not the rationale behind it. Methods such as Inverse Constitutional AI (ICAI) attempt to improve interpretability by compressing datasets into short ``constitutions'' of natural-language principles. We argue this framing is under-specified: a flat list of principles is not yet an executable decision rule because it leaves principle composition implicit. We use the pairwise setting as a testbed to empirically characterize three open problems in constitutional methods. First, principle quality is hard to measure: coverage and accuracy are useful but incomplete proxies for end-to-end reconstruction. Second, \emph{composition is ambiguous}: holding principles fixed, different executors (LLM judge versus majority vote) agree only $73\%$ of the time. Third, \emph{constitutions differ between LLMs}: cross-model vote agreement is $73\%$, whereas intra-model agreement is $81\%$. Across PRISM, AlpacaEval, and Chatbot Arena, we show that principle refinement (ICAI+) may be a first step towards ameliorating these problems: inter-executor agreement rises to $78\%$, and transparent executors match LLM judge accuracy ($66\%$ vs.\ $67\%$). Our results highlight that constitutions should be evaluated as \emph{constitution--executor systems}, with implications for LLMs-as-a-judge broadly.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.