붕괴 직전의 한 글자: 지시-튜닝된 유용성의 취약성
One Token Away from Collapse: The Fragility of Instruction-Tuned Helpfulness
지시-튜닝된 대규모 언어 모델은 유용하고 체계적인 답변을 생성하지만, 사소한 제약 조건이 가해질 때 이러한 유용성은 얼마나 견고할까요? 우리는 간단한 어휘적 제약(단일 구두 기호 또는 일반적인 단어 금지)이 지시-튜닝된 LLM의 답변을 붕괴시키고, 세 가지 공개 가중 모델 패밀리와 하나의 폐쇄 가중 모델(GPT-4o-mini)에 대한 쌍대 비교 평가에서 14~48%의 포괄성을 잃게 만든다는 것을 보여줍니다. GPT-4o-mini 및 GPT-4o가 판단한 1,920건의 쌍대 비교에서 기준 답변이 77~100%의 비율로 선호되었습니다. 주목할 점은 GPT-4o-mini가 31%의 포괄성 손실(99%의 기준 답변 우세율)을 겪었다는 점으로, 이는 서식 수준 제약에 대한 이전 연구 결과와 달리, 상용으로 배포된 폐쇄 가중 모델에도 이러한 취약성이 존재한다는 것을 보여줍니다. 메커니즘 분석을 통해, 이는 계획 실패로 인한 것임을 확인했습니다. 두 단계 생성(제약 없는 자유 생성 후 제약된 재작성)이 답변 길이를 59~96% 회복시켰습니다. 또한 프롬프트 표현에 대한 선형 프로브는 생성 시작 전에 답변 길이를 $R^2 = 0.51$에서 $0.93$ 사이의 정확도로 예측하며, 이 $R^2$ 값이 모델 간 붕괴 심각도를 추적합니다. 동일한 프로브는 기본 모델에서 음의 $R^2$ 값을 나타내어, 지시 튜닝이 붕괴 결정을 인코딩하는 표현 구조를 생성한다는 것을 확인합니다. 중요한 점은 기본 모델이 동일한 제약 조건 하에서 체계적인 붕괴를 보이지 않으며, 효과는 작고, 노이즈가 많으며, 양방향적이라는 것입니다. 이는 지시 튜닝이 작업 능력을 좁은 표면 형식 템플릿과 결합함으로써 이러한 취약성을 생성한다는 것을 보여줍니다. 이러한 효과는 MT-Bench의 모든 8가지 작업 범주에서 재현되었습니다. 또한 표준 독립적인 LLM-as-judge 평가 방법은 평균 3.5%의 품질 저하만 감지하는 반면, 쌍대 비교 평가에서는 23%의 품질 저하를 확인하여, 제약된 생성 평가에 대한 방법론적 한계점을 드러냅니다.
Instruction-tuned large language models produce helpful, structured responses, but how robust is this helpfulness when trivially constrained? We show that simple lexical constraints (banning a single punctuation character or common word) cause instruction-tuned LLMs to collapse their responses, losing 14--48% of comprehensiveness in pairwise evaluation across three open-weight model families and one closed-weight model (GPT-4o-mini). The baseline response is preferred in 77--100% of 1,920 pairwise comparisons judged by GPT-4o-mini and GPT-4o. Notably, GPT-4o-mini suffers 31% comprehensiveness loss (99% baseline win rate), demonstrating that the fragility extends to commercially deployed closed-weight models, contrary to prior findings on format-level constraints. Through mechanistic analysis, we identify this as a planning failure: two-pass generation (free generation followed by constrained rewriting) recovers 59--96% of response length, and linear probes on prompt representations predict response length with $R^2 = 0.51$--$0.93$ before generation begins, with $R^2$ tracking collapse severity across models. The same probes yield negative $R^2$ on base models, confirming that instruction tuning creates the representational structure encoding the collapse decision. Crucially, base models show no systematic collapse under identical constraints, with effects that are small, noisy, and bidirectional, demonstrating that instruction tuning creates this fragility by coupling task competence to narrow surface-form templates. The effect replicates on MT-Bench across all eight task categories. We further show that standard independent LLM-as-judge evaluation detects only a 3.5% average quality drop where pairwise evaluation reveals 23%, exposing a methodological blind spot in how constrained generation is assessed.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.