2602.00092v1 Jan 23, 2026 cs.LG

원자적 개념 편집을 위한 헌법을 활용한 모델 동작 해석 및 제어

Interpreting and Controlling Model Behavior via Constitutions for Atomic Concept Edits

Mani Malek
Mani Malek
Citations: 1,274
h-index: 6
N. Kalibhat
N. Kalibhat
Citations: 156
h-index: 5
Zi Wang
Zi Wang
Citations: 3,302
h-index: 4
Prasoon Bajpai
Prasoon Bajpai
Citations: 24
h-index: 4
Drew Proud
Drew Proud
Citations: 182
h-index: 1
Wenjun Zeng
Wenjun Zeng
Citations: 3,475
h-index: 5
Been Kim
Been Kim
Citations: 60
h-index: 2

본 논문에서는 검증 가능한 헌법을 학습하는 블랙박스 해석 가능성 프레임워크를 소개합니다. 이 헌법은 프롬프트의 변경이 모델의 특정 동작(예: 정렬성, 정확성 또는 제약 조건 준수)에 미치는 영향에 대한 자연어 요약입니다. 우리의 방법은 원자적 개념 편집(ACE, Atomic Concept Edits)을 활용합니다. ACE는 입력 프롬프트 내의 해석 가능한 개념을 추가, 제거 또는 대체하는 것을 목표로 하는 운영입니다. 다양한 작업에서 ACE를 체계적으로 적용하고, 그 결과 모델 동작에 미치는 영향을 관찰함으로써, 우리의 프레임워크는 편집과 예측 가능한 결과 사이의 인과 관계를 학습합니다. 이렇게 학습된 헌법은 모델에 대한 깊고 일반화 가능한 통찰력을 제공합니다. 실험적으로, 우리는 수학적 추론 및 텍스트-이미지 정렬을 포함한 다양한 작업에서 우리의 접근 방식을 검증하여 모델 동작을 제어하고 이해하는 데 사용했습니다. 텍스트-이미지 생성의 경우, GPT-Image는 문법 준수에 더 집중하는 경향이 있는 반면, Imagen 4는 대기적 일관성을 우선시합니다. 수학적 추론의 경우, 방해 변수는 GPT-5를 혼란스럽게 하지만 Gemini 2.5 모델과 o4-mini에는 큰 영향을 미치지 않습니다. 또한, 우리의 결과는 학습된 헌법이 모델 동작을 제어하는 데 매우 효과적이며, 헌법을 사용하지 않는 방법보다 평균적으로 1.86배 더 높은 성공률을 달성한다는 것을 보여줍니다.

Original Abstract

We introduce a black-box interpretability framework that learns a verifiable constitution: a natural language summary of how changes to a prompt affect a model's specific behavior, such as its alignment, correctness, or adherence to constraints. Our method leverages atomic concept edits (ACEs), which are targeted operations that add, remove, or replace an interpretable concept in the input prompt. By systematically applying ACEs and observing the resulting effects on model behavior across various tasks, our framework learns a causal mapping from edits to predictable outcomes. This learned constitution provides deep, generalizable insights into the model. Empirically, we validate our approach across diverse tasks, including mathematical reasoning and text-to-image alignment, for controlling and understanding model behavior. We found that for text-to-image generation, GPT-Image tends to focus on grammatical adherence, while Imagen 4 prioritizes atmospheric coherence. In mathematical reasoning, distractor variables confuse GPT-5 but leave Gemini 2.5 models and o4-mini largely unaffected. Moreover, our results show that the learned constitutions are highly effective for controlling model behavior, achieving an average of 1.86 times boost in success rate over methods that do not use constitutions.

2 Citations
0 Influential
3 Altmetric
17.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!