2607.14635v1 Jul 16, 2026 cs.AI

액션 Q포머: 비전-언어-액션 모델에서 액션 감독 하에 구조화된 표현 형성

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models

K. Sreenath
K. Sreenath
Citations: 13,589
h-index: 55
Zhongyu Li
Zhongyu Li
Citations: 2,033
h-index: 24
Yufeng Ji
Yufeng Ji
Citations: 5
h-index: 1
Wenhao Tang
Wenhao Tang
Tsinghua University
Citations: 78
h-index: 5
Haoyi Niu
Haoyi Niu
Tsinghua University
Citations: 277
h-index: 10
Yi Wu
Yi Wu
Citations: 76
h-index: 3

비전-언어-액션(VLA) 모델에서 액션 감독은 종종 액션 예측 학습을 위한 downstream 목표로 간주됩니다. 본 논문에서는 이를 오히려 상속된 다중 모드 표현을 형성하는 힘으로 연구합니다. 우리는 이러한 형성이 이중적인 효과를 가진다는 것을 보여줍니다. 즉, 액션과 호환되는 표현을 형성하는 데 필수적이지만, 액션 감독이 상속된 다중 모드 경로에 직접적으로 적용되면 언어 측 처리 및 객체 기반 연결을 지원하는 표현을 불안정하게 만들 수 있습니다. 이러한 긴장을 해결하기 위해 우리는 Action QFormer를 제안합니다. 이는 지시 사항에 조건화된 쿼리를 사용하여 상속된 다중 모드 정보를 액션 생성 전에 액션 중심의 표현으로 재구성하는 쿼리 기반 액션 인터페이스입니다. 제로샷 시뮬레이션-실제 내비게이션에서 Action QFormer는 평균 폐루프 작업 성공률을 18.8%에서 56.3%로 향상시키고, 고정된 지시 사항에 따른 액션 생성 정확도를 22.5%에서 75.5%로 높이며, 분포 외 지시 사항 생성을 거의 없앱니다. 추가 분석 결과, Action QFormer는 액션 감독이 상속된 다중 모드 표현을 형성하는 방식을 변화시켜 광범위한 상위 레벨 재작성을 줄이는 동시에 특정하고 때로는 긍정적인 액션 감독 하의 적응을 유지합니다. 이러한 결과는 VLA 성능 향상에 강력한 사전 학습 모델뿐만 아니라 상속된 다중 모드 정보를 선택하고 구성하는 더 나은 방법과 함께 액션 감독 하에서 정보가 어떻게 형성되는지를 제어하는 것이 중요하다는 것을 시사합니다.

Original Abstract

Action supervision in vision-language-action (VLA) models is often treated as a downstream objective for learning action prediction. In this paper, we study it instead as a force that shapes inherited multimodal representations. We show that this shaping has a dual effect: it is necessary for forming action-compatible representations, but when action supervision is applied too directly to the inherited multimodal pathway, it can also destabilize representations that support language-side processing and object grounding. To address this tension, we introduce Action QFormer, a query-based action-facing interface that uses instruction-conditioned queries to reorganize inherited multimodal information into action-facing representations before downstream action generation. In zero-shot sim-to-real navigation, Action QFormer improves average closed-loop task success from 18.8% to 56.3%, raises fixed-instruction action-generation correctness from 22.5% to 75.5%, and nearly eliminates out-of-distribution instruction generations. Further analyses show that Action QFormer changes how action supervision shapes inherited multimodal representations, reducing broad upstream rewriting while preserving targeted and sometimes constructive action-supervised adaptation. These results suggest that improving VLA performance requires not only stronger pretrained backbones, but also better ways of selecting and organizing inherited multimodal information while controlling how it is shaped under action supervision.

0 Citations
0 Influential
27.5 Altmetric
137.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!