DevicesWorld: 이기종 환경에서의 교차 장치 에이전트 성능 평가
DevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments
LLM 기반 에이전트는 모바일 애플리케이션, 데스크톱 시스템 및 스마트 홈과 같은 개별 디지털 환경에서 작동하는 능력이 빠르게 향상되었습니다. 그러나 실제 사용자의 목표는 종종 여러 장치를 포괄합니다. 예를 들어, 정보는 휴대폰에서 가져오고, 데스크톱에서 처리되며, 결과가 다른 장치에 표시되어야 할 수 있습니다. 대부분의 기존 벤치마크는 단일 환경에 중점을 두기 때문에 에이전트가 이기종 장치 간에 정보를 습득하고 통합하여 교차 장치 의존성을 갖는 엔드투엔드 작업을 수행할 수 있는지 평가하기 어렵습니다. 본 논문에서는 교차 장치 협업 운영을 위한 대규모 실행 가능한 벤치마크인 DevicesWorld를 소개합니다. DevicesWorld에는 6,140개의 작업이 포함되어 있으며, 모바일, 데스크톱 및 IoT의 세 가지 유형의 장치 환경을 통합하여 단일 교차 장치 상호 작용 및 평가 프레임워크를 제공합니다. 각 작업은 자연어 사용자 목표, 참여 장치 및 초기 상태, 실행 가능한 액션, 규칙 기반 검증기 및 정리 절차를 정의합니다. 다단계 구성 및 품질 관리 파이프라인을 통해 작업이 실제 사용자의 요구 사항과 가깝도록 유지하면서 최종 결과는 장치 상태 및 생성된 파일을 통해 자동으로 검증됩니다. 본 논문에서는 다섯 가지 최첨단 LLM 에이전트 시스템을 고정된 평가 세트를 사용하여 평가했습니다. 모든 방법은 낮은 성공률을 보였으며, 가장 좋은 성능을 보이는 모델조차도 12.5%에 불과했습니다. 실패한 실행 결과 중 약 28.7%는 적어도 하나의 점수 조건을 만족했지만 전체 작업에는 실패했습니다. 분석 결과, 에이전트가 정보를 습득하거나 인터페이스를 조작하는 데 어려움을 겪거나, 소스 장치와 출력 장치를 혼동하거나, 모든 조건이 동시에 충족되기 전에 종료되는 경우가 있었습니다. DevicesWorld는 신뢰할 수 있는 교차 장치 에이전트에 대한 연구를 위한 실행 가능하고 재현 가능한 진단적 평가 문제로 교차 장치 협업 운영을 제시합니다.
LLM-based agents have rapidly improved at operating individual digital environments such as mobile applications, desktop systems, and smart homes. However, real-world user goals often span multiple devices: information may come from a phone, be processed on a desktop, and the result may need to appear on another device. Most existing benchmarks center on a single dominant execution environment, making it difficult to evaluate whether agents can acquire and integrate information across heterogeneous devices and complete end-to-end tasks with cross-device dependencies. We introduce DevicesWorld, a large-scale executable benchmark for cross-device collaborative operation. DevicesWorld contains 6,140 tasks and integrates three classes of device environments -- mobile, desktop, and IoT -- into a unified cross-device interaction and evaluation framework. Each task defines a natural-language user goal, participating devices and initial states, executable actions, rule-based verifiers, and a cleanup procedure. A multi-stage construction and quality-control pipeline keeps tasks close to realistic user needs while allowing final outcomes to be automatically verified from device states and generated files. We evaluate five frontier LLM-agent systems on a fixed evaluation set. All methods achieve low success rates, with the best reaching only 12.5%. Among failed runs, about 28.7% satisfy at least one scoring condition yet still fail the full task. Trajectories show that agents become stuck acquiring information or manipulating interfaces, confuse source and output devices, or terminate before all conditions are jointly satisfied. DevicesWorld turns cross-device collaborative operation into an executable, reproducible, and diagnostically useful evaluation problem for research on reliable cross-device agents.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.