2605.26086v1 May 25, 2026 cs.AI

Claw-Anything: 사용자의 디지털 세계에 대한 보다 넓은 접근성을 갖춘 상시(Always-On) 개인 비서용 벤치마크

Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital World

Feiyang Pan
Feiyang Pan
Citations: 37
h-index: 4
Dandan Tu
Dandan Tu
Citations: 50
h-index: 3
Yusong Lin
Yusong Lin
Citations: 36
h-index: 3
Haiyang Wang
Haiyang Wang
Citations: 20
h-index: 2
Shuzhe Wu
Shuzhe Wu
Citations: 18
h-index: 2
Lu Fan
Lu Fan
Citations: 0
h-index: 0
Xinyu Liang
Xinyu Liang
Citations: 20
h-index: 3
Qi Gu
Qi Gu
Citations: 51
h-index: 5
Siqi Cheng
Siqi Cheng
Citations: 3
h-index: 1
Jiangui Chen
Jiangui Chen
Citations: 395
h-index: 8
Sanyuan Zhao
Sanyuan Zhao
Citations: 1,438
h-index: 14

거대 언어 모델 에이전트는 사용자 디지털 환경의 모든 관련 정보를 접근할 수 있는 상시 개인 비서로 점점 더 인식되고 있습니다. 그러나 현재 시스템은 그 세계의 아주 좁은 부분만 다루므로, 상황에 맞는 추론과 효과적인 지원을 제공하는 데 한계가 있습니다. 기존 벤치마크 또한 사용자 상태를 단편적으로만 제공하므로 이러한 광범위한 상시 환경에서 성능을 제대로 평가하지 못합니다. 이러한 공백을 메우기 위해, 본 연구에서는 에이전트의 컨텍스트를 세 가지 차원에서 확장합니다: 긴 호흡의 활동 내역, 상호 의존적인 백엔드 서비스, 그리고 여러 기기에서 통합된 GUI 및 CLI 인터페이스 간의 상호작용. 이러한 환경을 구현하기 위해, 우리는 다중 회차 이벤트 주입을 통해 수개월간의 사용자 활동을 시뮬레이션하여 복잡한 세계 상태와 실제적인 노이즈(무관한 이벤트 및 상충되는 신호)를 포함하게 합니다. 에이전트들은 풍부한 컨텍스트 환경에서 추론하며 이러한 노이즈에 대해 견고함을 유지해야 합니다. 이렇게 확장된 범위는 또한 에이전트가 사용자의 필요를 예측하고 시기적절한 권장 사항을 제시하도록 요구하는 선제적 지원(proactive assistance)의 평가도 가능하게 합니다. 실험 결과, GPT-5.5는 34.5%의 pass@1 성능을 달성했으며 이는 기존 벤치마크에서 훨씬 낮으며, 현재 에이전트 능력과 상시 개인 비서에 필요한 요구사항 사이의 간극을 강조합니다. 벤치마크와 함께, 우리는 2,000개의 학습 환경을 생성하는 자동화된 데이터 생성 파이프라인을 출시하며, 이를 통해 기반 모델을 23.7% 개선했으며, 이는 확장 가능한 데이터 인프라의 유용성을 입증합니다.

Original Abstract

Large language model agents are increasingly envisioned as always-on personal assistants with access to anything relevant in the user's digital world. Yet current systems operate over only narrow slices of that world, limiting context-sensitive reasoning and effective assistance. Existing benchmarks similarly provide only partial user state and therefore fail to capture performance in such a broad, always-on setting. To address this gap, we introduce Claw-Anything, a benchmark that expands agent context along three dimensions: long-horizon activity histories, interdependent backend services, and integrated GUI and CLI interaction across multiple devices. To instantiate this setting, we simulate months of user activity through multi-round event injection, producing complex world states and realistic noise, including irrelevant events and conflicting signals. Agents must reason over rich contextual environments while remaining robust to such noise. This expanded scope also enables the evaluation of proactive assistance, requiring agents to anticipate user needs and deliver timely recommendations. Experiments show that GPT-5.5 achieves only 34.5% pass@1, substantially below prior benchmarks, underscoring a gap between current agent capabilities and the demands of always-on personal assistance. Alongside the benchmark, we release an automated data-generation pipeline that yields 2,000 training environments and improves the base model by 23.7%, demonstrating its utility of scalable data infrastructure.

0 Citations
0 Influential
7 Altmetric
35.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!