FineVLA: 정밀한 지시 정렬을 통한 제어 가능한 시각-언어-행동 정책
FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies
시각-언어-행동(VLA) 모델은 로봇 작업을 수행하는 것뿐만 아니라, 사용자가 제시하는 작업 실행 방법에 대한 지시를 따르도록 점점 더 많이 요구되고 있습니다. 그러나 기존의 로봇 데이터셋은 일반적으로 궤적과 거친 목표 수준 언어를 연결하는데, 이는 활성 팔, 접근 방향 및 접촉 영역과 같이 작업 실행에 중요한 세부 정보를 명시하지 않습니다. 이로 인해 제어 가능한 정책 학습과 로봇 비디오 이해가 제한됩니다. 본 논문에서는 행동 정렬을 기반으로 한 정밀한 VLA 감독 방법을 제공하는 오픈 프레임워크인 FineVLA를 소개합니다. 이 프레임워크는 다음과 같은 구성 요소를 포함합니다: (1) 972,247개의 궤적을 85,000개의 작업에서 추출하여 10개의 오픈 소스 로봇 데이터셋을 통합하고, 인간이 검증한 47,159개의 정밀한 궤적으로 구성된 FineVLA-Data를 구축하는 데이터 생성 도구; (2) 500개의 비디오, 10,816개의 원자적 사실, 그리고 1,030개의 VQA 질문으로 구성된 별도의 벤치마크 세트; (3) 확장 가능한 정밀한 주석을 위한 로봇 전문 VLM 주석 도구; (4) 정밀한 수준의 지시와 일반적인 목표 수준 지시를 혼합하여 학습된 제어 가능한 VLA 정책. 실험 결과, 세 가지 중요한 사실이 밝혀졌습니다. 첫째, 정밀한 감독은 목표 수준의 성공률을 저해하지 않습니다. FineVaria만 사용했을 때 Raw-only 방식보다 +1.4에서 +8.1까지의 성공률 향상을 보였습니다. 둘째, 정밀한 지시와 일반적인 지시는 상호 보완적이며, FG:Raw 비율이 1:2에서 1:1 사이일 때 가장 높은 성능을 보입니다. 최적의 혼합 설정에서는 RoboTwin 시뮬레이션에서 86.8%/82.5%, 실제 양팔 조작 환경에서는 62.7/100의 성능을 달성했습니다 (Raw-only 방식은 각각 49.9). 셋째, 정밀한 감독은 제어 가능성을 향상시킵니다. 특히 자세 (+23), 색상 (+18) 및 접근 방향 (+18)과 같이 목표 수준의 지시로는 안내가 불가능한 요소에서 실제 환경에서의 성능 향상이 두드러졌습니다. 결론적으로, 정밀한 언어는 목표 수준의 지시를 보완해야 합니다. 즉, 달성해야 할 목표와 함께 실행 방법을 명확하게 제시해야 합니다. 프로젝트 페이지: https://finevla.xlang.ai/
Vision-Language-Action (VLA) models are increasingly expected to not only complete robot tasks, but also follow human instructions about how those tasks should be executed. However, existing robot datasets usually pair trajectories with coarse goal-level language, leaving execution-critical details such as active arm, approach direction, and contact region unspecified. This limits steerable policy learning and robotic video understanding. We introduce FineVLA, an open framework for action-aligned fine-grained VLA supervision. The framework includes: (1) a data construction tool that unifies 972,247 trajectories across 85K tasks from 10 open-source robot datasets and builds FineVLA-Data, a human-verified dataset of 47,159 fine-grained trajectories; (2) a held-out benchmark with 500 videos, 10,816 atomic facts, and 1,030 VQA questions; (3) a robotics-specialized VLM annotator for scalable fine-grained annotation; and (4) a steerable VLA policy trained with controlled mixtures of fine-grained and raw goal-level instructions. Our experiments yield three findings. First, fine-grained supervision does not sacrifice goal-level success: FG-only improves over Raw-only by +1.4 to +8.1 success-rate points across settings. Second, fine-grained and raw instructions are complementary, following a consistent inverted-U trend peaking at FG:Raw = 1:2 to 1:1. The best mixed setting reaches 86.8%/82.5% in RoboTwin simulation and 62.7/100 in real-world dual-arm manipulation (vs. 49.9 Raw-only). Third, fine-grained supervision improves steerable control: the largest real-world gains appear on pose (+23), color (+18), and approach direction (+18)--factors where goal-level instructions provide no guidance. Overall, fine-grained language should augment goal-level instructions: specifying how to execute alongside what to achieve. Project page: https://finevla.xlang.ai/
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.