2605.27354v1 May 26, 2026 cs.LG

희소 오토인코더 모델 내부 정보를 활용한 LLM 추가 학습 데이터 엔지니어링 가이드

Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders

Jinwu Hu
Jinwu Hu
Citations: 83
h-index: 5
Lei Hou
Lei Hou
Citations: 452
h-index: 8
Juanzi Li
Juanzi Li
Citations: 932
h-index: 13
Xiaozhi Wang
Xiaozhi Wang
Citations: 108
h-index: 4
Yi Jing
Yi Jing
Citations: 4
h-index: 1
Zao Dai
Zao Dai
Citations: 0
h-index: 0
Zijun Yao
Zijun Yao
Tsinghua University
Citations: 682
h-index: 12

모델의 내부 정보는 대규모 언어 모델(LLM)이 학습 데이터를 어떻게 처리하는지에 대한 풍부한 정보를 담고 있지만, 추가 학습 데이터 엔지니어링은 주로 외부 신호에 의존하며 모델 내부에 존재하는 중요한 내부 신호를 간과합니다. 본 논문에서는 LLM 강화 학습(RL)을 위한 데이터 엔지니어링 프레임워크인 SAERL을 제안합니다. SAERL은 희소 오토인코더(SAE)를 사용하여 추출된 모델 내부 정보로부터 다양성, 난이도 및 품질의 세 가지 내재적 데이터 속성을 모델링합니다. 각 속성은 구체적인 데이터 엔지니어링 연산에 기반하며, SAE-공간 클러스터링을 통한 적절한 배치 혼합으로 배치 다양성을 제어하고, 난이도 지표를 활용하여 쉬운 것부터 어려운 것 순서로 학습 과정을 구성하며, 품질 검사를 통해 데이터를 필터링합니다. SAERL은 Qwen2.5-Math-1.5B 모델에서 기존 GRPO 알고리즘보다 평균 정확도를 3.00% 향상시키고, 목표 정확도에 도달하는 데 필요한 학습 단계를 20% 줄였습니다. 이러한 성능 향상은 다양한 모델 크기와 RL 알고리즘에서도 일관적으로 나타났습니다. 실험 결과는 SAE가 다양한 모델 계열과 규모에서 효과적으로 활용될 수 있으며, 가볍고 재사용 가능한 데이터 엔지니어링 도구로 기능할 수 있음을 보여줍니다. 이러한 결과는 모델 내부 정보가 추가 학습 데이터 엔지니어링을 위한 강력하고 실용적인 정보원임을 입증합니다.

Original Abstract

Model internals encode rich information about how a large language model (LLM) processes its training data; however, post-training data engineering largely relies on external signals and ignores rich intrinsic signals lying in model internals. We propose SAERL, a data engineering framework for LLM reinforcement learning (RL). It models three intrinsic data properties: diversity, difficulty, and quality, using model internals extracted with Sparse Autoencoder (SAE), an advanced mechanistic interpretability tool. Each property grounds a concrete data engineering operation: SAE-space clustering with moderate batch mixing for batch diversity control, a difficulty proxy for easy-to-hard curriculum ordering, and a quality probe for data filtering. SAERL improves average accuracy by 3.00% over vanilla GRPO and reaches target accuracy with 20% fewer training steps on Qwen2.5-Math-1.5B, with consistent gains across model scales and RL algorithms. Experiments show that SAE transfers effectively across model families and scales, serving as a lightweight and reusable data engineering tool. These results demonstrate that model internals are a powerful and practical source of signals for post-training data engineering.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!