2608.05144v1 Aug 05, 2026 cs.AI

Argus: 장기적인 추론을 위한 범용 에이전트 기반 실행 환경

Argus: A General-Purpose Agentic Reasoning Runtime for Long-Horizon Tasks

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Yijia Fan
Yijia Fan
Citations: 440
h-index: 12
Xuanhe Zhou
Xuanhe Zhou
Citations: 14
h-index: 2
Chuan Wen
Chuan Wen
Citations: 344
h-index: 5
Zimo Wen
Zimo Wen
Citations: 5
h-index: 1
Boxiu Li
Boxiu Li
Citations: 27
h-index: 3
Yifei Shen
Yifei Shen
Citations: 27
h-index: 3
Xian Zhang
Xian Zhang
Citations: 178
h-index: 5
Junxiang Lei
Junxiang Lei
Citations: 0
h-index: 0
Ruize Tang
Ruize Tang
Citations: 25
h-index: 2
Xiaoyu Chen
Xiaoyu Chen
Citations: 0
h-index: 0
R. Gu
R. Gu
Citations: 0
h-index: 0
Zhijie Deng
Zhijie Deng
Citations: 74
h-index: 4
Sufeng Guo
Sufeng Guo
Citations: 0
h-index: 0
Jiaao Wu
Jiaao Wu
Citations: 129
h-index: 1
Mukai Li
Mukai Li
Citations: 13
h-index: 2
Wanbo Zhang
Wanbo Zhang
Citations: 22
h-index: 2
Yifei Gao
Yifei Gao
Citations: 15
h-index: 3
Xuyao Huang
Xuyao Huang
Citations: 17
h-index: 2
Zelong Zhao
Zelong Zhao
Citations: 0
h-index: 0
Shi-Wei Hu
Shi-Wei Hu
Citations: 0
h-index: 0
Hang Guo
Hang Guo
Citations: 0
h-index: 0
Yilin Chen
Yilin Chen
Citations: 0
h-index: 0
Yuzhe Zhang
Yuzhe Zhang
Citations: 0
h-index: 0
Fan Yang
Fan Yang
Citations: 56
h-index: 4

장기적인 추론은 현재 접근 방식에 대한 증거가 있을 때는 지속하고, 실패, 숨겨진 제약 조건 또는 잘못된 목표가 드러날 때에는 전환할 수 있는 에이전트 기반의 실행 환경을 필요로 합니다. 본 논문에서는 관리자, 계획자, 엔지니어 및 검토자가 영구적인 프로젝트 상태에서 제한적인 작업을 수행하는 지속적이고 자체 진화하는 실행 환경인 Argus를 소개합니다. Argus는 안정적인 사용자 의도를 운영 목표, 제약 조건 및 검증 기준으로 분리하며, 역할을 기반으로 한 검토와, 가능하다면 작업에 내재된 검증을 거쳐야만 메모리, 기술, 절차, 검증기, 라우팅 결정 및 거부된 경로가 허용됩니다. 모델 가중치는 고정되어 있으며, 자체 진화는 영구적인 런타임 상태와 제어 정책을 통해 이루어지며, 운영자가 관리하는 에스컬레이션 지점 사이에 자율적으로 실행됩니다. Argus는 GPT-5.5 벤치마크 환경에서 SWE-Bench Pro에서 약 78%의 정확도를 달성했으며, 이는 Direct Copilot의 59%보다 높은 수치입니다 (총 토큰 사용량은 1.41배). 검증을 거친 자체 진화를 통해, 성숙된 SWE-Bench 단계는 초기 단계에 비해 문제 해결 입력 토큰을 21% 적게 사용하고 활성 워크플로우 시간을 작업당 평균 15% 절약했습니다. 또한 34건의 검증기 복구 및 22건의 엄격한 검토 루프를 통한 문제를 해결했습니다. Argus는 AARRI-Bench에서 76.8%의 정확도를 달성했으며, 수학 데이터 합성 분야에서는 28.0점의 격차를 보이며, GPU 커널 및 언어 모델 학습에서도 경쟁력 있는 결과를 보여주었습니다. 벤치마크 외에도 최적화된 RWKV6 커널이 상위 레벨로 통합되었으며, 다일간의 수학 캠페인에서는 잘못된 경로를 유지하고 증거 기반의 프론티어 업데이트를 수행했으며, 6개의 논문 파이프라인에서 16번의 단계 되돌리기를 포함하여 총 254건의 작업을 완료했습니다. 이러한 결과는 고정 가중치와 자체 진화 기능을 갖춘 시스템이 검증된 접근 방식을 수정, 복구 및 축적하면서 향후 지도 학습 및 강화 학습을 위한 체계적인 경로를 생성할 수 있음을 보여줍니다.

Original Abstract

Long-horizon reasoning requires an agentic runtime that can persist when evidence supports its current approach and pivot when measurements reveal failure, hidden constraints, or a misspecified objective. We present Argus, a persistent, self-evolving runtime in which Manager, Planner, Engineer, and Reviewer execute bounded missions over durable project state. Argus separates stable user intent from operational objectives, constraints, and verification criteria, and admits memories, skills, procedures, verifiers, routing decisions, and rejected routes only after role-owned review and, when available, task-native verification. Model weights remain fixed; self-evolution occurs through persistent runtime state and control policy, with autonomous execution between operator-owned escalation points. Across seven GPT-5.5 benchmark arenas, Argus achieves about 78% on SWE-Bench Pro versus 59% for Direct Copilot while using 1.41 times the aggregate tokens. After verification-gated self-evolution, mature SWE-Bench waves use 21% fewer solve-input tokens and 15% less active workflow time per task than startup waves, while recording 34 verifier recoveries and 22 strict review-loop rescues. Argus also reaches 76.8% on AARRI-Bench and a 28.0-point gap on mathematical data synthesis, with competitive GPU-kernel and language-model-training results. Beyond benchmarks, an optimized RWKV6 kernel was merged upstream; a multi-day mathematics campaign retained falsified routes and proof-backed frontier updates; and six paper pipelines completed 254 missions with 16 stage rollbacks. These results show that a fixed-weight, self-evolving harness can revise, recover, and accumulate verified approaches while producing structured trajectories for future supervised and reinforcement learning.

0 Citations
0 Influential
6 Altmetric
30.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!