2606.19750v1 Jun 18, 2026 cs.LG

다양체 방랑자: 대규모 언어 모델의 잠재 기하학 구조를 활용한 베이지안 커리큘럼 학습

Manifold Bandits: Bayesian Curriculum Learning over the Latent Geometry of Large Language Models

Xiaolong Wang
Xiaolong Wang
Citations: 2,101
h-index: 10
Nicklas Hansen
Nicklas Hansen
Citations: 3,740
h-index: 19
Darrien M. McKenzie
Darrien M. McKenzie
Citations: 2
h-index: 1

강화 학습(RL)은 대규모 언어 모델(LLM)의 추론 능력을 향상시키는 핵심적인 방법이며, 학습 효율성은 최적화 과정에서 문제 샘플링 방식에 크게 의존합니다. 기존의 적응형 커리큘럼 학습 방법은 일반적으로 중간 난이도의 프롬프트를 우선시하며, 문제 선택을 독립적인 臂를 가진 표준 방랑자 문제로 간주하여 작업 공간의 구조적이고 이질적인 특성을 간과합니다. 본 연구에서는 문제 샘플링을 모델의 잠재 표현 공간을 통해 연결된 내생적 비정상성을 갖는 다양체 구조의 방랑자 문제로 정의합니다. 즉, 샘플링 결정은 학습 신호가 해당 공간에서 어떻게 진화하는지에 영향을 미칠 수 있습니다. 이러한 관점을 실현하기 위해, 우리는 구조를 고려한 프레임워크인 베이지안 다양체 커리큘럼(BMC)을 소개합니다. BMC는 문제를 계층적 작업 트리로 구성하고, 베이지안 학습을 사용하여 샘플링을 안내합니다. 실험적으로, 다양한 샘플링 전략이 생산성(학습 신호), 다양성(작업 다양체의 범위), 유용성(평가 관련성) 간에 중요한 상충 관계를 초래한다는 것을 확인했습니다. 이러한 결과는 단순히 난이도를 우선시하는 것만으로는 강력한 성능을 달성하기에 충분하지 않으며, 문제 샘플링에 구조와 유형 인지 기능을 통합하는 것이 중요하다는 점을 강조합니다.

Original Abstract

Reinforcement learning (RL) is a central approach for improving reasoning capabilities in large language models (LLMs), where training efficiency depends critically on how problems are sampled during optimization. Existing adaptive curriculum learning methods typically prioritize prompts of intermediate difficulty, treating problem selection as a standard bandit problem with independent arms and overlooking the structured, heterogeneous nature of the task space. In this work, we frame problem sampling as a manifold-structured bandit problem with endogenous non-stationarity: problems are related through the model's latent representation space, and sampling decisions can steer how learning signals evolve across that space. To operationalize this perspective, we introduce Bayesian Manifold Curriculum (BMC), a structure-aware framework that organizes problems into a hierarchical task tree and applies Bayesian learning to guide sampling. Empirically, we find that different sampling strategies induce non-trivial tradeoffs between productivity (learning signal), diversity (coverage of the task manifold), and utility (evaluation relevance). These results show that prioritizing difficulty alone is insufficient for strong downstream performance, highlighting the importance of incorporating structure and type-awareness into problem sampling.

0 Citations
0 Influential
9.5 Altmetric
47.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!