Qwen-CUA: (거의) 모든 것에 활용 가능한 네이티브 컴퓨터 사용
Qwen-CUA: Native Computer Use for (almost) Everything
네이티브 컴퓨터 사용은 에이전트가 사람이 사용할 수 있는 거의 모든 소프트웨어를 작동할 수 있도록 하는 일반적인 인터페이스를 제공하지만, 장기적인 상태 추적, 대규모의 상호 작용 경험, 그리고 희소하지만 검증 가능한 결과로부터 학습하는 것이 필요합니다. 본 논문에서는 397B-A17B Qwen Mixture-of-Experts 아키텍처를 기반으로 하는 네이티브 컴퓨터 사용 에이전트인 Qwen-CUA를 소개합니다. Qwen-CUA는 DOM 트리, 접근성 메타데이터 또는 작업별 API 없이 화면 캡처만 관찰하고 키보드 및 마우스 이벤트를 통해 작동합니다. 시스템은 최대 20개의 활성 화면 캡처를 유지하며, 오래된 시각적 정보를 고정 크기의 블록으로 저장하여 최근 증거를 보존하면서 재사용 가능한 프롬프트 접두사를 유지합니다. 학습을 위해, 거의 10만 개의 vCPU와 수만 개의 동시 환경에 접근할 수 있는 클라우드 환경을 구축하고, 약 4만 개의 검증 가능한 작업을 구성했으며, 일상 및 전문 소프트웨어 전반에 걸쳐 개인화된 장기 워크플로우를 수집했습니다. 시스템은 검증 가능한 보상을 사용하여 전체 경로를 최적화하며, 경로 분할을 통해 학습합니다. 반복적인 학습 과정에서 지도 데이터가 업데이트되고 강화 학습 작업을 재조정합니다. Qwen-CUA는 8개의 벤치마크에서 Qwen3.7보다 우수한 성능을 보이며, 선도적인 독점 시스템과 경쟁력 있는 성능을 보여주며, OSWorld-Verified에서는 86.2, OSWorld 2.0의 이진/부분 완료율은 각각 18.5/48.4를 달성했습니다. 동일한 방식을 1조 개 이상의 파라미터를 가진 모델에 적용하면 Qwen-CUA-Max가 생성되며, 이는 해당 점수를 각각 87.6 및 21.2/53.3으로 향상시킵니다. 또한 Qwen-CUA는 Qwen3.7과 비교하여 RedTeamCUA 공격 성공률을 36.6에서 16.4로 감소시켰습니다. 효율성 분석, 브라우저 배포 및 Bash 기반 실험을 통해 실제 동작을 추가적으로 분석했습니다. 이러한 결과는 네이티브 컴퓨터 사용이 광범위하게 활용 가능한 에이전트의 기반이 될 수 있으며, 검증 가능한 상호 작용과 하이브리드 도구 사용이 중요한 발전 방향임을 보여줍니다.
Native computer use offers a general interface for agents to operate almost any software available to people, but requires long-horizon state tracking, large-scale interactive experience, and learning from sparse yet verifiable outcomes. We introduce Qwen-CUA, a native computer-use agent with a 397B-A17B Qwen mixture-of-experts backbone. It observes only screenshots and acts through keyboard and mouse events, without DOM trees, accessibility metadata, or task-specific APIs. Its scaffold maintains up to 20 active screenshots and folds older visual history in fixed-size blocks to retain recent evidence while preserving reusable prompt prefixes. For training, we build a cloud rollout fleet with access to nearly 100,000 vCPUs and tens of thousands of concurrent environments, construct approximately 40,000 verifiable tasks, and collect personalized long-horizon workflows across everyday and professional software. We optimize complete trajectories with verifiable rewards and trajectory slicing, while iterative training runs refresh supervised data and recalibrate reinforcement-learning tasks. Across eight benchmarks, Qwen-CUA outperforms Qwen3.7 and remains competitive with leading proprietary systems, reaching 86.2 on OSWorld-Verified and 18.5/48.4 binary/partial completion on OSWorld 2.0. Scaling the same recipe to a model with over one trillion parameters yields Qwen-CUA-Max, improving these scores to 87.6 and 21.2/53.3. Qwen-CUA also reduces RedTeamCUA attack success from 36.6 to 16.4 relative to Qwen3.7. Efficiency analyses, a browser deployment, and Bash-augmented experiments further characterize practical behavior. These results establish native computer use as a broadly capable agent foundation and highlight scalable verifiable interaction and hybrid tool use as key directions.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.