2606.31524v1 Jun 30, 2026 cs.LG

자기 개선 온라인 LLM 정렬의 수렴성에 대한 연구

On the Convergence of Self-Improving Online LLM Alignment

Pangpang Liu
Pangpang Liu
Citations: 32
h-index: 3
Xudong Wu
Xudong Wu
Citations: 0
h-index: 0
Vaneet Aggarwal
Vaneet Aggarwal
Citations: 213
h-index: 7
Jiayu Chen
Jiayu Chen
Citations: 64
h-index: 5

자기 개선 정렬 (SAIL) 알고리즘은 데이터 분포 변화에 대처하기 위해 문제의 양층 구조를 효율적인 단일층 방법으로 줄입니다. 실험적으로, SAIL은 이 작업에서 뛰어난 성능을 보여주었습니다. 그러나 그 수렴 특성에 대한 공식적인 분석은 부족했습니다. 우리는 핵심적인 이론적 문제를 발견했는데, 이는 표준 SAIL 목적 함수가 Hessian의 불리한 특성으로 인해 강하게 오목(strongly concave)하다는 보장이 없다는 것입니다. 이러한 한계를 해결하기 위해, 우리는 역 Kullback-Leibler (KL) 발산 페널티를 통합하여 최적화 환경을 개선하는 정규화된 목적 함수인 SAIL-RevKL을 제안합니다. 우리의 주요 이론적 기여는 이 정규화된 목적 함수가 특정 경계 내의 파라미터 공간에서 Polyak-Lojasiewicz (PL) 조건을 만족한다는 것을 증명하는 것입니다. 우리는 전역 수렴 보장을 확립하여 거의 선형적인 샘플 복잡도를 달성했습니다. 또한, 우리는 실험적 평가를 통해 SAIL-RevKL의 효과성과 안정성을 검증했으며, MuJoCo 벤치마크 및 LLM 정렬 작업 모두에서 기존의 SAIL보다 우수한 성능을 보이는 것을 확인했습니다.

Original Abstract

The Self-Improving Alignment (SAIL) algorithm addresses distribution shift by reducing a bilevel formulation of the problem to an efficient, single-level method. Empirically, SAIL has demonstrated strong performance on this task. However, a formal analysis of its convergence properties has been lacking. We identify a key theoretical challenge: the standard SAIL objective function is not guaranteed to be strongly concave due to unfavorable properties of its Hessian. To address this limitation, we propose a regularized objective, SAIL-RevKL, which incorporates a reverse Kullback-Leibler (KL) divergence penalty to improve the optimization landscape. Our central theoretical contribution is to prove that this regularized objective satisfies the Polyak-Lojasiewicz (PL) condition within a bounded parameter space. We establish global convergence guarantees, achieving a near-linear sample complexity. We further validate the effectiveness and stability of SAIL-RevKL through empirical evaluations, demonstrating that it outperforms the vanilla SAIL on both MuJoCo benchmarks and LLM alignment tasks.

1 Citations
0 Influential
3.5 Altmetric
18.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!