2608.04934v1 Aug 05, 2026 cs.CL

State2State: 환경 기반의 LLM 에이전트 중간 학습

State2State: Environment-Derived Mid-Training for LLM Agents

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Chenliang Li
Chenliang Li
Citations: 76
h-index: 5
Xuanyu Lei
Xuanyu Lei
Citations: 100
h-index: 4
Ming Yan
Ming Yan
Citations: 849
h-index: 11
Yang Liu
Yang Liu
Citations: 339
h-index: 7

LLM 에이전트의 학습은 일반적으로 전문가의 경로를 이용한 지도 미세 조정 또는 인간이 정의한 작업에 대한 온라인 강화 학습을 통해 이루어집니다. 이러한 방법들은 효과적이지만, 외부에서 지정된 작업 및 감독 신호에 의해 제한되어 에이전트 학습의 확장성과 다양성을 저해합니다. 본 연구에서는 에이전트가 외부로 설정된 작업 없이 환경과의 상호 작용을 통해서만 상호 작용 및 조작 능력을 습득하는 환경 학습 패러다임을 제시합니다. 우리는 State2State라는 환경 기반 중간 학습 방법을 제안하며, 이는 탐색된 환경 상태를 학습 목표로 변환하여 에이전트에게 특정 대상 상태에 도달하도록 도전합니다. State2State는 환경 탐색으로부터 작업을 파생하고 규칙 기반의 상태 매칭을 통해 성공 여부를 검증함으로써, 전문가 감독이나 수동 작업 설계 없이 확장 가능하고 검증 가능한 학습 목표를 제공합니다. ALFWorld 및 ScienceWorld에서의 실험 결과, State2State는 대부분의 경우 독립적인 환경 학습 단계로서 에이전트 성능을 향상시킵니다. 또한, 다운스트림 강화 학습의 초기화로 사용될 때 최종 성능과 학습 효율성을 더욱 향상시키며, 다양한 환경에 대한 일반화 가능성이 있음을 보여줍니다.

Original Abstract

Training LLM agents commonly relies on supervised fine-tuning from expert trajectories or online reinforcement learning over human-specified tasks with handcrafted verifiers. Though effective, both remain bottlenecked by externally specified tasks and supervision signals, limiting the scalability and diversity of agent training. We study an environment learning paradigm in which agents acquire interaction and manipulation capabilities solely through environment interaction, without externally specified tasks. We propose State2State, an environment-derived mid-training method that converts explored environment states into training objectives, challenging agents to reach a specified target state. By deriving tasks from environment exploration and verifying success through rule-based state matching, State2State provides scalable and verifiable training objectives without expert supervision or manual task design. Experiments on ALFWorld and ScienceWorld show that State2State improves agent performance as a standalone environment-learning stage in most settings. As initialization for downstream RL, it further improves final performance and learning efficiency, with promising evidence of cross-environment generalization.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!