강화 학습: 알고리즘부터 기반 모델까지
Reinforcement Learning: From Algorithms To Foundation Models
강화 학습(RL)은 명시적인 목표 하에서 순차적 의사 결정을 위한 프레임워크를 제공합니다. 고전적인 형태의 강화 학습은 에이전트가 동적 환경에서 장기적인 보상을 극대화하기 위해 어떻게 행동해야 하는지를 연구합니다. 더욱 발전된 환경에서는 단일 에이전트와 고정된 환경을 넘어, 전략적 상호 작용, 불확실성에 대한 적응, 그리고 고차원 세계에 대한 추론이 지능적인 행동에 필요합니다. 본 논문은 강화 학습을 두 가지 관점에서 연구합니다: 게임에서의 알고리즘과 기반 모델 시대의 강화 학습. 첫 번째 부분에서는 게임 내 다중 에이전트 강화 학습에 초점을 맞춥니다. 이는 경쟁적 및 일반 합 환경에서 인센티브, 정책, 그리고 균형 개념이 어떻게 상호 작용하는지를 살펴보고, 두 플레이어 제로섬 게임, 대규모 비디오 게임, 그리고 일반적인 구조를 가진 멀티플레이어 설정을 포함합니다. 이러한 연구는 다중 에이전트 시스템에서의 학습과 상호 작용 환경에서 강화 학습 방법의 행동을 조사합니다. 두 번째 부분은 사전 지식이 순차적 의사 결정을 풍부하게 할 수 있다는 아이디어에 기반하여 생성 및 기반 모델과의 강화 학습을 연구합니다. 사전 훈련된 생성 모델과 학습된 세계 모델은 계획, 제어 및 정책 최적화를 위한 표현 도구이자 구조화된 사전 정보로 사용됩니다. 본 논문에서는 확산 기반의 세계 모델을 개발하고, 효율적인 비디오 생성을 위한 강화 학습을 조사하며, 생성 모델을 정책 클래스로 탐색하고, 행동이 미래 관찰에 영향을 미치는 상호 작용 비디오 세계 모델을 연구합니다. 또한 메모리를 갖는 아키텍처를 통해 장기적 모델링 문제를 해결합니다. 이러한 기여는 함께 강화 학습을 복잡한 순차적 영역에서의 목표 지향적인 적응이라는 통합된 관점에서 제시합니다. 전략적 게임부터 생성 세계 모델까지, 본 논문은 강화 학습이 의사 결정, 환경 모델링, 그리고 새롭게 등장하는 기반 모델의 기능을 어떻게 연결하는지를 강조하며, 지능적인 행동에 대한 더 넓은 관점을 제공합니다.
Reinforcement learning (RL) provides a framework for sequential decision making under explicit objectives. In its classical form, RL studies how an agent should act to maximise long-term reward in a dynamic environment. In richer settings, the problem extends beyond a single agent and fixed environment: intelligent behavior may require strategic interaction, adaptation to uncertainty, and reasoning over high-dimensional worlds. This thesis studies RL from two perspectives: algorithms in games and RL in the era of foundation models. The first part focuses on multi-agent RL in games. It examines how incentives, policies, and equilibrium concepts interact in competitive and general-sum environments, spanning two-player zero-sum games, large-scale video games, and multi-player settings with general structure. These works investigate learning in multi-agent systems and the behavior of RL methods in interactive environments. The second part studies RL with generative and foundation models, motivated by the idea that prior knowledge can enrich sequential decision making. Pretrained generative models and learned world models serve as representation tools and structured priors for planning, control, and policy optimization. The thesis develops diffusion-based world models, investigates RL for efficient video generation, explores generative models as policy classes, and studies interactive video world models in which actions shape future observations. It also addresses long-horizon modeling through architectures with memory. Together, these contributions present a unified view of RL as objective-driven adaptation in complex sequential domains. From strategic games to generative world models, the thesis highlights how RL connects decision making, environment modeling, and emerging foundation-model capabilities, offering a broader perspective on the principles underlying intelligent behavior.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.