2605.29267v1 May 28, 2026 cs.AI

인간의 검토가 실패하는 시점과 방법: 다중 모델 기반 자기 참조 루프 환경에서의 선호도 일치

When and How Human Curation Backfires: Preference Alignment under Multi-Model Self-Consuming Loop

Xueru Zhang
Xueru Zhang
Citations: 755
h-index: 15
Xiukun Wei
Xiukun Wei
Citations: 22
h-index: 2
Yang Zhang
Yang Zhang
Citations: 267
h-index: 5

최근 기초 모델은 이전 모델 반복에서 생성된 합성 데이터를 활용하여 훈련되는 경우가 많으며, 이는 실제 데이터만을 사용하는 것보다 일반적입니다. 이러한 자기 참조 학습 방식은 모델의 성능 저하, 발산 또는 편향 증폭을 초래할 수 있습니다. 최근 연구(Ferbach et al., 2024)에 따르면 인간의 검토를 학습 루프에 통합하면 자기 참조 모델을 인간의 선호도와 일치하도록 유도할 수 있지만, 이러한 분석은 주로 자체 출력만 사용하는 단일 모델에 집중합니다. 그러나 실제로는 모델들이 서로 상호 작용하며 다른 모델이 생성한 입력-출력 쌍으로 훈련되는 경우가 많습니다. 본 논문에서는 다중 모델 환경에서의 자기 참조 학습을 연구합니다. 먼저 상호 작용하는 자기 참조 모델을 위한 프레임워크를 정의하고, 결과적인 동적 시스템이 안정적인 상태로 수렴하는 조건을 분석합니다. 또한 하나의 모델에 대한 인간의 검토가 해당 모델 자체의 정렬(자기 영향)에 미치는 영향을 조사하고, 이러한 효과가 다른 모델에 어떻게 전파되는지(교차 영향) 분석합니다. 단일 모델 환경에서 인간의 검토가 항상 모델의 정렬을 향상시키는 것과는 달리, 다중 모델 간의 상호 작용은 이러한 효과를 약화시키거나 심지어 역전시켜 장기적인 정렬 성능을 저하시킬 수 있음을 보여줍니다.

Original Abstract

Foundation models are increasingly trained on synthetic data generated by prior model iterations rather than exclusively on real data. This self-consuming training paradigm can lead to model collapse, divergence, or bias amplification. Recent work (Ferbach et al., 2024) shows that incorporating human curation into the loop can steer a self-consuming model toward human-aligned behavior, but these analyses focus on a single, isolated model that solely consumes its own outputs. In practice, however, models often interact and train on input-output pairs produced by other models. This paper studies self-consuming training in the multi-model regime. We first formalize a framework for interacting self-consuming models and characterize when the resulting dynamical system converges to a stable point. We then examine how human curation of one model affects its own alignment (self-influence) and how such effects propagate to other models (cross-influence). Unlike isolated settings where human curation always enhances model alignment, we show that cross-model interactions can dampen or even invert this effect, ultimately degrading long-term alignment.

0 Citations
0 Influential
7.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!