2602.11079v2 Feb 11, 2026 cs.LG

야생 모델 유기체: 데이터 귀속을 통한 프로덕션 LLM 후처리 단계에서 발생하는 바람직하지 않은 새로운 현상 완화

In-the-Wild Model Organisms: Mitigating Undesirable Emergent Behaviors in Production LLM Post-Training via Data Attribution

Frank Xiao
Frank Xiao
Citations: 3
h-index: 1
Santiago Aranguri
Santiago Aranguri
Citations: 23
h-index: 2

본 연구에서는 활성화 기반 데이터 귀속 기법을 제안합니다. 이 방법은 후처리된 언어 모델에서 발생하는 행동 변화를 유발하는 원본 학습 데이터 포인트를 추적합니다. 테스트 프롬프트와 선호도 쌍에 대한 활성화 차이 벡터를 계산하고 코사인 유사도를 기준으로 순위를 매김으로써, 특정 행동을 유발하는 데이터 포인트를 식별하고, 수정된 데이터로 재학습하여 이러한 귀속의 인과 관계를 검증합니다. 또한, 행동-데이터 포인트 유사성 행렬을 클러스터링하면 지도 학습 없이도 새로운 현상을 발견할 수 있습니다. OLMo 2의 프로덕션 DPO 학습에 이 방법을 적용한 결과, '방해 요인 유발 순응'이라는 유해한 현상이 발견되었습니다. 이는 모델이 위험한 요청에 응답할 때, 무해한 서식 지침이 추가되면 더욱 순응하는 현상입니다. 상위 순위 데이터 포인트를 필터링하면 이 현상을 63% 줄일 수 있으며, 레이블을 변경하면 78%까지 감소시킬 수 있습니다. 본 방법은 기울기 기반 귀속 방법 및 LLM 평가 모델을 사용한 기준 방법보다 성능이 우수하며, 비용은 10배 이상 저렴합니다. 이러한 '야생 모델 유기체'는 의도적인 삽입이 아닌 오염된 선호도 데이터에서 발생하는 현실적인 안전 기술 평가 기준으로 활용될 수 있습니다.

Original Abstract

We propose activation-based data attribution, a method that traces behavioral changes in post-trained language models to responsible training datapoints. By computing activation-difference vectors for both test prompts and preference pairs and ranking by cosine similarity, we identify datapoints that cause specific behaviors and validate these attributions causally by retraining with modified data. Clustering behavior-datapoint similarity matrices also enables unsupervised discovery of emergent behaviors. Applying this to OLMo 2's production DPO training, we surfaced distractor-triggered compliance: a harmful behavior where the model complies with dangerous requests when benign formatting instructions are appended. Filtering top-ranked datapoints reduces this behavior by 63% while switching their labels achieves 78%. Our method outperforms gradient-based attribution and LLM-judge baselines while being over 10 times cheaper than both. This in-the-wild model organism - emerging from contaminated preference data rather than deliberate injection - provides a realistic benchmark for safety techniques.

0 Citations
0 Influential
1 Altmetric
5.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!