2606.16802v1 Jun 15, 2026 cs.AI

LabOSBench: 과학 기기 제어를 위한 컴퓨터 사용 에이전트 성능 평가

LabOSBench: Benchmarking Computer Use Agents for Scientific Instrument Control

Zhaoyang Liu
Zhaoyang Liu
Citations: 125
h-index: 4
Ben Fei
Ben Fei
Citations: 205
h-index: 9
Han Deng
Han Deng
Citations: 13
h-index: 2
Anqi Zou
Anqi Zou
Citations: 8
h-index: 2
Wanli Ouyang
Wanli Ouyang
Citations: 33
h-index: 2
Chengyun Zhang
Chengyun Zhang
Citations: 1
h-index: 1
Junquan Hu
Junquan Hu
Citations: 5
h-index: 1
Yu Wang
Yu Wang
Citations: 0
h-index: 0
Yuxiang Xing
Yuxiang Xing
Citations: 29
h-index: 2
Aokai Zhang
Aokai Zhang
Citations: 4
h-index: 2
Hanling Zhang
Hanling Zhang
Citations: 170
h-index: 5
Zhihui Wang
Zhihui Wang
Citations: 17
h-index: 3

현재의 컴퓨터 사용 성능 평가 도구는 주로 가상화된 시스템에서의 소프트웨어 운영 작업에 초점을 맞추고 있습니다. 반면, 과학 기기 환경에서는 복잡한 인터페이스에 대한 통합적인 제어와 피드백 기반 파라미터 조정이 필요합니다. 그러나 실제 고정밀 기기에 에이전트를 직접 평가하는 것은 높은 비용, 안전 문제, 제한된 접근성 및 재현 가능한 평가 보장 어려움 때문에 비현실적입니다. 따라서 본 연구에서는 과학 기기의 운영상의 어려움을 유지하면서도 확장 가능하고 안전한 성능 평가를 가능하게 하는 시뮬레이션 기반의 현실적인 테스트베드를 개발하고자 합니다. 이를 위해 웹 기반 과학 기기 시뮬레이터 모음으로 구축된, 다중 인터페이스 에이전트를 위한 도전적인 벤치마크인 LabOSBench를 소개합니다. LabOSBench는 리소스 집약적인 OS 가상화를 피하고 유연한 작업 구성 및 실행 기반 평가를 지원하며, 웹 브라우저를 통해 직접 작동합니다. 구체적으로, LabOSBench는 샘플 로딩, 정렬, 파라미터 조정, 데이터 획득부터 결과 검토에 이르는 워크플로우를 포함하는 8개의 기기 시뮬레이터를 사용하여 총 96개의 하위 작업으로 구성됩니다. 본 연구에서는 범용 비전-언어 모델, 특화된 GUI 에이전트 모델 및 고급 에이전트 프레임워크를 하위 작업 수준과 전체 워크플로우 수준 모두에서 평가합니다. 실험 결과, 기존 에이전트는 많은 구조화된 GUI 하위 작업을 완료할 수 있지만, 피드백 기반 운영 및 장기적인 워크플로우 실행에는 여전히 어려움을 겪는 것으로 나타났습니다. 전반적으로 LabOSBench는 과학 기기 제어를 위한 컴퓨터 사용 에이전트의 발전을 가속화하는 데 도움이 되는 재현 가능하고 저렴한 테스트베드를 제공합니다.

Original Abstract

Current computer-use benchmarks primarily focus on software operation tasks in virtualized systems, whereas scientific instrumentation scenarios require coordinated control over complex interfaces, and feedback-driven parameter adjustment. However, directly evaluating agents on physical high-precision instruments is impractical due to high cost, safety risks, limited accessibility, and difficulty in ensuring reproducible evaluation. This motivates the need for a simulated yet realistic testbed that preserves the operational challenges of scientific instruments while enabling scalable and safe benchmarking. To this end, we introduce LabOSBench, a challenging benchmark for multimodal GUI agents built on a suite of web-based scientific-instrument simulators. Operating directly via a browser, LabOSBench avoids resource-heavy OS virtualization while supporting flexible task configuration and execution-based evaluation. Specifically, LabOSBench constructs 96 subtasks across eight instrument simulators, covering workflows from sample loading, alignment, parameter tuning, and data acquisition to result inspection. We evaluate general-purpose vision-language models, specialized GUI agent models, and advanced agentic frameworks at both subtask and end-to-end levels. Our experiments reveal that while existing agents can complete many structured GUI subtasks, they still struggle with feedback-driven operations and long-horizon workflow execution. Overall, LabOSBench provides a reproducible, low-cost testbed for advancing computer-using agents toward scientific-instrument control.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!