2606.11182v1 Jun 09, 2026 cs.LG

EEVEE: 현실 세계에서의 자기 개선 에이전트를 위한 테스트 시간 프롬프트 학습

EEVEE: Towards Test-time Prompt Learning in the Real World for Self-Improving Agents

Shilong Liu
Shilong Liu
Citations: 426
h-index: 7
Mengdi Wang
Mengdi Wang
Citations: 745
h-index: 13
Weixia Xu
Weixia Xu
Citations: 38
h-index: 2

본 논문에서는 LLM 에이전트를 위한 최초의 다중 데이터셋 기반 테스트 시간 프롬프트 학습 프레임워크인 EEVEE를 제안합니다. EEVEE는 실제 환경에서의 작업 스트림 하에서 테스트 시간 프롬프트 학습을 가능하게 합니다. 기존 방법들은 대부분 단일 데이터셋 환경에 맞춰 설계되었으며, 실제 응용 분야에서는 모델이 여러 데이터셋, 도메인 및 작업 분포로부터 유입되는 이질적인 입력 스트림을 처리해야 하기 때문에 실용성이 제한됩니다. 이러한 데이터셋 간의 간섭 문제를 완화하기 위해 EEVEE는 수신된 입력을 작업 클러스터로 분할하고 적절한 프롬프트 구성으로 할당하는 라우터를 도입합니다. 이 설계는 라우터와 프롬프트 학습 단계를 번갈아 수행하는 라우터-프롬프트 공동 진화 전략을 통해 최적화됩니다. 여러 데이터셋에 대한 실험 결과, EEVEE는 이질적인 데이터 스트림 하에서 안정성을 향상시키는 동시에 개별 벤치마크 학습 능력과 효율성을 유지합니다. 특히, EEVEE는 Qwen3-4B-Instruct 및 DeepSeek-V3.2 모델에서 평균 다중 벤치마크 점수를 각각 10.38점 및 24.32점 향상시켰으며, SOTA 방법인 GEPA 및 ACE를 최대 37.2% 및 48.2%까지 능가했습니다.

Original Abstract

In this paper, we propose EEVEE, the first multi-dataset test-time prompt learning framework for LLM agents, enabling test-time prompt learning under real-world task streams. Existing methods are largely designed for single-dataset settings, while real-world applications require models to handle heterogeneous input streams drawn from multiple datasets, domains, and task distributions, limiting their practical applicability. To mitigate cross-dataset interference, EEVEE introduces a router that partitions incoming inputs into task clusters and assigns them to suitable prompt configurations. This design is optimized via a router-prompt co-evolution strategy, which employs interleaved router and prompt learning phases to address their mutual dependency. Experiments across multiple datasets demonstrate that the framework improves robustness under heterogeneous data streams while maintaining single-benchmark learning capability and efficiency. Specifically, EEVEE improves average multi-benchmark scores by 10.38 and 24.32 points over Qwen3-4B-Instruct and DeepSeek-V3.2, surpassing SOTA methods GEPA and ACE by up to 37.2% and 48.2%.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!