2606.20002v1 Jun 18, 2026 cs.LG

점 연결: 강화 학습을 통한 교차 도메인 일반화로 장기 수명 에이전트를 위한 LLM 훈련

Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning

Jingren Zhou
Jingren Zhou
Citations: 26,959
h-index: 29
Yuexiang Xie
Yuexiang Xie
Citations: 1,510
h-index: 20
Yaliang Li
Yaliang Li
Citations: 1,839
h-index: 19
Yanxi Chen
Yanxi Chen
Citations: 244
h-index: 8
Weijie Shi
Weijie Shi
Citations: 25
h-index: 3
Bolin Ding
Bolin Ding
Citations: 1,124
h-index: 14
Bo Hu
Bo Hu
Citations: 0
h-index: 0

본 연구는 대규모 언어 모델(LLM)을 훈련하여 "점 연결(Connect the Dots, CoD)"이라는 메타 역량을 갖추도록 하는 일반적인 프레임워크를 제시합니다. CoD는 장기 수명 에이전트에 필요한 핵심 기능으로, LLM 기반 AI 에이전트가 특정 환경에 배치되면, 해당 환경을 지속적으로 탐색하면서 일련의 작업을 해결하고, 자신의 경험으로부터 학습하며, 환경에 대한 자체적인 맥락을 반복적으로 업데이트하여 미래 작업에서 더 나은 성능을 달성합니다. CoD 프레임워크의 주요 구성 요소는 다음과 같습니다. (1) 긴 시퀀스 길이의 강화 학습(RL) 알고리즘 설계 및 인프라, 여기서 작업 해결과 컨텍스트 업데이트가 번갈아 가며 수행됩니다. (2) 훈련 중에 LLM에서 목표로 하는 메타 역량을 유도하고 장려하기 위한 작업 및 환경, 그리고 평가 과정에서 진행 상황을 정확하게 측정하기 위한 환경입니다. 본 연구에서는 GRPO 스타일의 RL 알고리즘과 세분화된 신용 할당 기능을 포함한 CoD 프레임워크의 개념 증명 구현체와, 목표 메타 역량에 맞춰 설계된 작업 및 환경(특정 도메인 LLM 기능이나 표준 작업별 RL이 아닌)을 제시합니다. 실험 결과는 CoD 설정을 통한 엔드투엔드 RL 훈련의 효과를 검증하고, 훈련 도메인 내, 서로 다른 도메인 간, 그리고 CoD에서 Ralph-loop 설정으로의 일반화 가능성을 보여줍니다. 본 연구는 기존 연구들의 다양한 측면을 연결하고, LLM 및 AI 에이전트 발전에 새로운 기회를 제공합니다. 추가적인 연구 및 응용을 위해, 저희 구현체를 다음 URL에서 공개합니다: https://github.com/agentscope-ai/Trinity-RFT/tree/research/cod/examples/research_cod

Original Abstract

This work presents a general framework for training large language models (LLMs) to "Connect the Dots" (CoD), a meta-capability required by long-lifecycle agents: as an LLM-based AI agent gets deployed in an environment, it solves a long sequence of tasks while continuously exploring the environment, learning from its own experiences, and iteratively self-updating its context about the environment, thereby achieving progressively better performance on future tasks conditioned on the updated context. Major components of the CoD framework include: (1) algorithm design and infrastructure for end-to-end reinforcement learning (RL) with long rollout sequences interleaving solve-task and update-context episodes; (2) tasks and environments for incentivizing and eliciting the targeted meta-capability in LLMs during training, as well as for faithfully measuring progress during evaluation. We present proof-of-concept implementations of the CoD framework, including a GRPO-style RL algorithm with fine-grained credit assignment, as well as tasks and environments tailored to the targeted meta-capability (rather than domain-specific LLM capabilities or standard task-by-task RL). Empirical results validate the efficacy of end-to-end RL training in the CoD setting, and demonstrate the potential for out-of-distribution generalization -- within the training domains, across different domains, and from CoD to Ralph-loop settings -- of the elicited meta-capability. Our investigation of CoD connects several lines of prior works, and opens up new opportunities for advancing LLMs and AI agents. To facilitate further research and applications, we release our implementations at \url{https://github.com/agentscope-ai/Trinity-RFT/tree/research/cod/examples/research_cod}.

0 Citations
0 Influential
66.923176178176 Altmetric
0.0 Score
Original PDF
654

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!