LabVLA: 과학 실험실 환경에서의 비전-언어-행동 모델 적용
LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories
과학 실험실은 점점 더 많은 AI 시스템에 의존하여 실험을 수행하지만, 실제 과학적 활동은 아직까지 AI의 한계 내에 있습니다. AI는 문헌 분석, 가설 생성 및 프로토콜 계획을 지원할 수 있지만, 이러한 프로토콜의 실행은 여전히 인간 작업자의 개입이 필요합니다. 비전-언어-행동(VLA) 모델은 서면 프로토콜과 로봇 실행 간의 가능한 인터페이스를 제공하지만, 기존 모델들은 주로 가정 환경 및 탁상 시연 데이터로 학습되어 과학 실험실에서 흔히 발견되는 장비, 투명한 액체 또는 고정된 프로토콜 워크플로우에 대한 경험이 부족합니다. 이러한 격차를 해소하기 위해서는 실험실 특정 데이터와 다양한 로봇 플랫폼을 수용할 수 있는 통합 학습 프레임워크가 필요합니다. 따라서 우리는 데이터 및 로봇 플랫폼(embodiment)을 핵심적인 문제점으로 파악했습니다. 데이터 측면에서, 구성된 실험실 워크플로우를 기본 기술 단위로 조합하고, 실행 과정을 검증 및 필터링하며, 다양한 로봇 프로필에 맞는 구조화된 시연 데이터를 생성하는 시뮬레이션 기반 워크플로우 및 데이터 엔진인 RoboGenesis를 개발했습니다. 정책 측면에서, 우리는 두 단계 학습법으로 훈련된 LabVLA 모델을 제시합니다. 먼저, FAST 액션 토큰 사전 학습을 통해 Qwen3-VL-4B-Instruct 백본에 액션 인지 능력을 부여하고, 이후 Flow Matching 후속 학습을 통해 지식 격리 환경에서 DiT 액션 전문가를 연결합니다. LabUtopia 벤치마크 테스트 결과, LabVLA는 평가된 모든 기본 모델 중에서 in-distribution 및 out-of-distribution 환경 모두에서 가장 높은 평균 성공률을 달성했습니다.
Scientific laboratories increasingly rely on AI systems to reason about experiments, but the physical act of doing science remains largely outside their reach. AI can help read literature, generate hypotheses, and plan protocols, yet the execution of those protocols at the bench still requires a human operator. Vision-Language-Action (VLA) models provide one possible interface between written protocols and robot execution, but existing policies are trained mostly on household and tabletop demonstrations and rarely encounter the instruments, transparent liquids, or fixed protocol workflows found in scientific laboratories. Closing this gap requires both laboratory-specific supervision and a unified learning framework that can accommodate the diverse robot embodiments used to execute experimental protocols. We therefore identify data and embodiment as central bottlenecks alongside model design. To address the data side, we build RoboGenesis, a simulation-based workflow and data engine that composes configured laboratory workflows from atomic skills, validates and filters rollouts, and exports structured demonstrations across supported robot profiles. On the policy side, we present LabVLA, trained with a two-stage recipe: FAST action token pretraining first makes the Qwen3-VL-4B-Instruct backbone action aware before any continuous control is learned, and flow matching posttraining then attaches a DiT action expert under knowledge insulation. On the LabUtopia benchmark, LabVLA achieves the highest average success rate among all evaluated baselines under both in-distribution and out-of-distribution settings.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.