2608.05013v1 Aug 04, 2026 cs.CL

OneDayAgent: 자율 에이전트를 위한 장기적인 활용 프레임워크

OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents

Ningyu Zhang
Ningyu Zhang
Citations: 333
h-index: 9
Jingsheng Zheng
Jingsheng Zheng
Citations: 65
h-index: 3
Jintian Zhang
Jintian Zhang
Citations: 847
h-index: 11
Zhengke Gui
Zhengke Gui
Citations: 171
h-index: 5
Huajun Chen
Huajun Chen
Citations: 208
h-index: 5
Xinyuan Fang
Xinyuan Fang
Citations: 0
h-index: 0

LLM 기반 에이전트는 업무, 학습, 생활 등 다양한 분야의 복잡한 요청을 처리하는 데 점점 더 많이 활용되고 있습니다. 이러한 작업은 장기간에 걸쳐 여러 환경과 멀티모달 데이터를 다루며, 에이전트는 목표와 제약 조건을 유지하면서 다양한 도구와 리소스를 효율적으로 사용해야 합니다. 기존 연구에서는 목표 편향, 상태 손실, 컨텍스트 오버플로우 등 개별적인 문제점에 대한 해결책을 제시했지만, 단일 프레임워크가 이러한 문제점들을 종합적으로 관리하고 다양한 백엔드에서 효과를 유지하는지에 대한 연구는 부족했습니다. 본 논문에서는 자율 에이전트를 위한 장기 활용 프레임워크인 OneDayAgent를 제안합니다. OneDayAgent는 복잡한 요청을 관리 가능한 실행 프로세스로 변환하여, 작업을 제한된 하위 작업으로 분해하고, 컨텍스트 압박 상황에서도 실행 메모리를 유지하며, 최종 결과물을 검증하고 수정하는 기능을 제공합니다. 우리는 AgentIF-OneDay 데이터셋의 104개 작업에 대해 OneDayAgent를 평가했습니다. GLM-5.2 백엔드를 사용했을 때, OneDayAgent는 전체 점수 0.821으로 최고 성능을 달성했습니다. 또한, 동일한 프레임워크가 세 가지 모델 패밀리의 다섯 가지 백엔드 LLM에서 실행 가능했으며, 이는 별도의 튜닝 없이도 다양한 백엔드에서 일반화될 수 있음을 나타냅니다. 각 모델은 동일한 워크플로우 하에서도 고유한 실행 방식을 보여줍니다.

Original Abstract

LLM agents are increasingly applied to open-ended everyday requests that span work, study, and life. These tasks are long-horizon, cross-environment, and multimodal, forcing the agent to preserve goals and constraints across many steps while navigating heterogeneous tools and attachments. While prior work has addressed individual failure modes such as goals drift, states loss, and context overflow, whether a single harness can manage them jointly and remain effective across backends has received less study. We present OneDayAgent, a long-horizon harness for autonomous agents. OneDayAgent turns an open-ended request into a managed execution process that decomposes tasks into bounded subtasks, maintains execution memory under context pressure, and verifies and repairs the final deliverable. We evaluate OneDayAgent on AgentIF-OneDay across 104 tasks. With the GLM-5.2 backend, OneDayAgent sets a new state of the art with an overall score of 0.821. The same harness runs across five backend LLMs from three model families, indicating the harness generalizes across backends without tuning, even as different models induce distinct execution styles under the same workflow.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!