역방향 강화 학습이 인간 모방을 통해 AI 정렬에 기여한다
Inverse RL Helps Align AI by Imitating Humans
언어 모델 정렬은 모델의 동작이 유용성, 안전성 및 지시 따르기와 같은 바람직한 특성을 안정적으로 반영하도록 하는 것을 목표로 합니다. 현재의 접근 방식은 일반적으로 데모를 사용한 지도 미세 조정 또는 검증기나 인간 피드백으로부터 파생된 보상을 사용하는 강화 학습을 활용합니다. 이러한 패러다임은 중요한 질문 하나를 충분히 탐구하지 못합니다: 데모만으로 AI 정렬을 위한 암묵적인 보상을 얻어내고, 이를 검사하고 재사용하며 정책에 맞춰 최적화할 수 있을까요? 역방향 강화 학습에 영감을 받아, 저희는 데모로부터 추정된 투영 정렬 보상(Projected Alignment Reward Estimated from Demonstrations, PARED)을 소개합니다. PARED는 전문가의 데모에 내재된 암묵적인 보상을 작은 응답 수준 기능 집합에 대한 명시적인 함수로 복원합니다. 이는 경량 디스크리미네이터가 이 기능 공간에서 데모와 정책 자체의 샘플을 구분하도록 학습하여 이루어집니다. 표준적인 보상 모델과 달리, PARED는 작업별 선호도 주석이 필요하지 않습니다. 데모는 작업별 감독 신호를 제공하며, 이는 추가적인 감독 차원으로 AI 피드백을 통해 강화될 수 있습니다. 추론 시간 재정렬 및 적대적 온-정책 강화 학습 실험을 통해, 복원된 보상이 지도 손실 없이 기본 정책을 개선하고 표준적인 지도 미세 조정 후에 최적화될 때 더 큰 성능 향상을 가져옴을 보여줍니다. 또한, PARED가 컨텍스트 정렬에 사용될 수 있음을 입증했습니다. 이를 통해 하나의 정책을 다양한 청중의 선호도에 맞게 조정할 수 있습니다.
Language model alignment aims to make model behavior reliably reflect desirable properties such as helpfulness, safety, and instruction following. Current approaches typically use supervised fine-tuning on demonstrations or reinforcement learning with rewards derived from verifiers or human feedback. These paradigms leave an important question underexplored: can demonstrations alone yield an implicit reward that can be inspected, reused, and optimized on-policy to align AI? Motivated by inverse reinforcement learning, we introduce Projected Alignment Reward Estimated from Demonstrations (PARED). PARED recovers the implicit reward underlying expert demonstrations as an explicit function over a small set of response-level features, learned by a lightweight discriminator that separates demonstrations from the policy's own samples in this feature space. Unlike a standard reward model, PARED requires no task-specific preference annotations: demonstrations provide the task-specific supervision, which can be augmented with AI feedback as additional dimensions of supervision. Through experiments involving inference-time reranking and adversarial on-policy RL, we show that the recovered reward improves a base policy without a supervised loss and yields further gains when optimized after standard supervised fine-tuning. Additionally, we demonstrate that PARED can be used for contextual alignment, in which a single policy can be tailored to the preferences of different audiences.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.