2605.29888v1 May 28, 2026 cs.LG

LaRA: 계층별 표현 분석을 통한 강화 학습 추가 훈련 과정에서의 데이터 오염 탐지

LaRA: Layer-wise Representation Analysis for Detecting Data Contamination in RL Post-Training

Alan Ritter
Alan Ritter
Citations: 9
h-index: 2
Minju Gwak
Minju Gwak
Yonsei University
Citations: 127
h-index: 3
Minseok Kwak
Minseok Kwak
Citations: 2
h-index: 1
Dongseok Lee
Dongseok Lee
Citations: 21
h-index: 2
Guijin Son
Guijin Son
Citations: 3
h-index: 1
Jaehyung Kim
Jaehyung Kim
Citations: 3
h-index: 1

강화 학습(RL) 기반의 추가 훈련은 대규모 언어 모델(LLM)의 추론 능력을 향상시키는 것으로 나타났습니다. 그러나, RL 추가 훈련 과정에서 발생하는 데이터 오염 문제는 일반화 성능을 저해하고 훈련 과정 자체의 신뢰성을 떨어뜨릴 수 있음에도 불구하고, 이에 대한 연구는 미흡한 실정입니다. 기존의 탐지 방법은 주로 likelihood(확률)나 entropy(엔트로피)와 같은 출력 수준의 신호에 의존하는데, 이는 RL 모델이 토큰 확률보다는 trajectory-level(경로 수준)의 보상을 통해 행동을 학습하기 때문에 신뢰성이 떨어집니다. 본 연구에서는 RL 추가 훈련된 LLM에서 데이터 오염을 탐지하기 위한 계층별 표현 분석 프레임워크인 LaRA를 제안합니다. LaRA는 perturbation sensitivity(perturbation에 대한 민감도), directional collapse(방향성 함몰), local representation rigidity(지역적 표현의 강직성)를 측정하는 세 가지 상호 보완적인 지표를 도입하여, 통제된 perturbation 하에서 모델을 분석합니다. 연구 결과, 데이터 오염은 계층 전반에 걸쳐 점진적인 기하학적 왜곡을 유발하며, 이는 perturbation sensitivity 증폭, directional collapse 심화, local rigidity 강화로 나타납니다. 이러한 발견을 바탕으로, 우리는 표현 수준의 왜곡을 계층 및 지표별로 종합하는 데이터 오염 탐지 프로토콜을 개발했습니다. RL 기반 추론 모델에 대한 실험 결과, 제안하는 프로토콜은 기존의 출력 수준 기반 방법보다 데이터 오염 탐지에 더 우수한 성능을 보였습니다.

Original Abstract

Reinforcement learning (RL) post-training has shown to improve reasoning in large language models (LLMs). However, there has been little exploration on the problem of data contamination in RL post-training, potentially undermining generalization and evaluation reliability of the training process itself. Existing detection methods primarily rely on output-level signals such as likelihood or entropy, which become unreliable for RL-trained models since RL shapes behavior through trajectory-level rewards rather than token likelihoods. We propose LaRA, a layer-wise representation analysis framework for detecting contamination in RL post-trained LLMs. LaRA introduces three complementary metrics, measuring perturbation sensitivity, directional collapse, and local representation rigidity under controlled perturbations. We find that contamination produces progressive geometric deviations across layers, including amplified perturbation sensitivity, stronger directional collapse, and enhanced local rigidity. Based on our findings, we also develop a contamination detection protocol that aggregates representation-level deviations across layers and metrics. Experiments on RL-trained reasoning models show that our protocol outperforms existing output-level baselines for contamination detection.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!