OSReward: 크로스 플랫폼 컴퓨터 사용 보상 모델에 대한 표준화된 평가 시스템 구축
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
컴퓨터 사용 에이전트(CUA)는 디지털 세상 전반에서 빠르게 발전하고 있습니다. CUA의 동작 경로는 에이전트의 행동, 상태 및 추론 과정을 기록합니다. 작업 지시를 충족했는지 확인하는 것은 CUA 평가, 데이터 큐레이션 및 강화 학습에 매우 중요합니다. 인간이 작성한 검증 도구나 주석가는 대규모로 이러한 검증을 제공할 수 없기 때문에, 이 분야에서는 점점 더 많은 연구가 CUA 경로의 판단자로 사용되는 비전-언어 모델(VLM)에 의존하고 있습니다. 그러나 오랫동안 간과되어 온 근본적인 질문이 있습니다: 이러한 VLM 판단자는 충분히 신뢰할 만한가? 이를 체계적으로 연구하기 위해, 우리는 현실적이고 고품질의 벤치마크인 OSReward를 소개합니다. 이 벤치마크는 다양한 에이전트 아키텍처에서 실행되는 인간 검증된 작업 지시를 기반으로 하며, 여러 단계를 거치는 인간 주석을 통해 정확한 결과를 갖도록 설계되었습니다. 이를 바탕으로, 우리는 특히 어려운 사례에 집중된 OSReward-Hard 및 세밀한 효율성 및 정렬 점수를 위한 OSReward-Multi 데이터셋을 개발했습니다. 현재까지 수행된 VLM 판단자에 대한 가장 포괄적인 평가 결과, 최첨단 모델조차 이상적인 판단자 수준에는 미치지 못하며, 실패한 실행 결과를 성공으로 잘못 분류하는 체계적인 관대함 편향이 있음을 확인했습니다. 신뢰할 만하다고 여겨지는 모델은 비용이 너무 많이 들고, 저렴한 공개 모델은 성능이 훨씬 뒤쳐집니다. 이러한 격차를 해소하기 위해, 우리는 CUA 커뮤니티를 위한 추론 주석이 포함된 경로 판단 데이터셋인 OS-Shepherd-100K를 구축하고 공개했습니다. 이 데이터셋을 사용하여 OS-Shepherd (9B 및 35B)이라는 오픈 소스 보상 모델을 학습시켰으며, 이는 상용 판단자와 비교하여 30~60% 더 낮은 비용으로 안정적이고 신뢰할 수 있는 보상 신호를 제공합니다. 추가적인 분석은 대규모 CUA 보상의 설계에 대한 중요한 정보를 제공합니다. 저희의 코드, 벤치마크, 데이터셋 및 모델 체크포인트는 https://os-copilot.github.io/OSReward-Home/ 에서 확인할 수 있습니다.
Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the task instruction is central to CUA evaluation, data curation, and reinforcement learning. Neither human-written verifiers nor human annotators can provide such verification at scale, so the field increasingly turns to vision-language models (VLMs) as judges of CUA trajectories. But a fundamental question has long gone unexamined: are these VLM judges reliable enough? To study it systematically, we introduce OSReward, a realistic, high-quality benchmark that evaluates VLM judges on CUA trajectories. The trajectories come from diverse agent backbones executing human-verified instructions across platforms, and are then rigorously labeled with ground-truth verdicts through multi-stage human annotation. Building on it, we derive OSReward-Hard, a challenge set concentrating genuinely hard cases, and OSReward-Multi for fine-grained efficiency and alignment scoring. The most comprehensive evaluation of VLM judges to date finds even state-of-the-art models fall short of an ideal judge, sharing a systematic leniency bias that mislabels failed runs as successes. The few reliable enough to trust are too expensive to run at scale, while affordable open models trail far behind. To close this gap, we construct and release OS-Shepherd-100K, an open corpus of reasoning-annotated trajectory judgments for the CUA community. On it, we train OS-Shepherd (9B and 35B), open reward models that supply low-cost, stable, and reliable reward signals, matching commercial judges at 30-60x lower cost than the frontier. Extensive analyses further inform the design of reliable CUA reward at scale. Our code, benchmark, dataset, and model checkpoints are available at https://os-copilot.github.io/OSReward-Home/.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.