ODYSSE: 에피소드 기반 정책 최적화를 통한 개인화된 에이전트 추론
ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning
에이전트 시스템은 실제 환경과의 상호작용 능력, 외부 도구 활용 능력, 그리고 사용자에게 서비스 제공 능력에서 급속한 발전을 이루었습니다. 하지만 자연 세계의 작업처럼 명확하게 정의된 지침을 가정하는 것과는 달리, 인간 중심 시나리오는 모호한 요청으로 인해 광범위하고 개방적인 해결 공간을 특징으로 합니다. 따라서 사용자의 개인화된 선호도를 파악하여 후보 솔루션 범위를 좁히는 것이 필수적입니다. 이는 새로운 과제인 개인화된 에이전트 추론으로 이어지는데, 이 과정에서 에이전트는 사용자 및 환경과 함께 상호 작용하며 개인화된 서비스를 제공해야 합니다. 본 논문에서는 개인화된 에이전트 추론을 위한 강화 학습 기반 미세 조정(RFT) 프레임워크인 ODYSSE를 제시합니다. 핵심적으로, ODYSSE는 에피소드별 그룹 상대 정책 최적화(ESPO)라는 새로운 방법을 제안하는데, 이는 장기적인 행동 범위와 개인화된 에이전트 추론에서 발생하는 강력한 단계 간 의존성을 해결하기 위해 고안된 Group Relative Policy Optimization (GRPO)의 확장입니다. ESPO는 개별 단계를 독립적으로 최적화하는 대신, 에피소드 수준의 보상 메커니즘과 에피소드 기반 이점 추정 기능을 도입하여 상위 계층의 증거가 하위 계층의 개인화된 의사 결정을 효과적으로 안내하고, 에이전트가 여러 단계의 상호 작용을 통해 모호한 사용자 요청을 점진적으로 해결할 수 있도록 합니다. 또한, 동일 에피소드의 행동을 통합된 학습 배치로 묶는 에피소드 기반 배치 샘플러를 제안하여 ESPO 하에서 일관성 있는 최적화를 지원합니다. ODYSSE는 현실적인 장기 개인화 GUI 추론 작업에 대해 평가되었으며, 실험 결과는 ODYSSE가 전문 분야 및 범용 LVLM 모두보다 우수한 성능을 보임을 보여주며, 이는 개인화된 에이전트 추론에 대한 효과성을 강조합니다.
Agentic systems have rapidly advanced in their ability to interact with real-world environments, leverage external tools, and provide services for users. However, unlike natural-world tasks that assume well-defined instructions, human-centered scenarios are characterized by ambiguous requests that lead to large, open-ended solution spaces. Decoding users' personalized preferences is therefore essential for narrowing the candidate solution space. This introduces a new challenge, personalized agentic reasoning, which requires agents to jointly interact with both users and environments to deliver personalized services. In this paper, we present ODYSSE, a Reinforced Fine-Tuning (RFT) framework for personalized agentic reasoning. At its core, ODYSSE proposes Episode-wise GRPO (ESPO), a novel extension of Group Relative Policy Optimization (GRPO) designed to address long action horizons and strong cross-step dependencies in personalized agentic reasoning. Rather than optimizing individual steps independently, ESPO introduces an episode-level reward mechanism together with episodic advantage estimation, enabling upstream evidence to effectively guide downstream personalized decisions and allowing agents to progressively resolve ambiguous user requests across multiple interaction steps. We further propose an episodic batch sampler that groups actions from the same episode into unified training batches, facilitating coherent optimization under ESPO. We evaluate ODYSSE on realistic long-horizon personalized GUI reasoning tasks. Experimental results demonstrate that ODYSSE consistently outperforms both specialist and general-purpose LVLMs, highlighting its effectiveness for personalized agentic reasoning.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.