2607.10744v4 Jul 12, 2026 cs.CV

Traj-VLN: 자기 회귀 경로 생성을 통한 픽셀 공간 상호작용 학습

Traj-VLN: Learning Pixel-Space Interaction via Autoregressive Trajectory Generation

Hong Zhang
Hong Zhang
Citations: 66
h-index: 5
Changfei Fu
Changfei Fu
Citations: 37
h-index: 3
Guangcheng Chen
Guangcheng Chen
Citations: 39
h-index: 4
Aoxiang Gu
Aoxiang Gu
Citations: 5
h-index: 1
Haoxiang Liang
Haoxiang Liang
Citations: 0
h-index: 0
Wenjun Xu
Wenjun Xu
Citations: 17
h-index: 1

대규모 사전 훈련 데이터에 내재된 강력한 사전 지식과 새롭게 등장하는 상식 추론 능력을 바탕으로, 대규모 언어 모델(LLM)은 다양한 연구 분야에서 전례 없는 일반화 능력을 보여주고 있습니다. 최근에는 비전-언어 모델(VLM)을 통해 시각적 임베딩을 언어 공간에 투영하여 시뮬레이션 환경과 실제 환경, 그리고 다양한 장면에서의 일반화를 달성하는 것이 연속적인 환경에서의 비전-언어 탐색(VLN-CE) 분야의 주요 패러다임이 되었습니다. VLN은 에이전트가 자연스러운 언어 지시를 따라 아직 보지 못한 환경을 탐색하도록 하는 것을 요구합니다. 본 연구에서는 VLN 작업을 일련의 하위 작업으로 분해할 수 있으며, 각 하위 작업은 "소파 끝까지 걸어가서 왼쪽으로 방향을 틀어라"와 같은 지시에 의해 설명되는 3차원 공간 상호작용 프로세스에 해당한다고 강조합니다. 그러나 깊이 감지를 통해 이미지 내에서 움직이는 것과 관련된 이러한 공간 상호작용은 주로 RGB 이미지를 사용한 대화로 훈련된 VLM에게는 어려운 과제입니다. 따라서 본 연구에서는 깊이나 3차원 기하학적 정보를 통합하는 대신(VLM이 사전 훈련 중에 거의 접하지 못하는 정보), VLM을 미세 조정하여 자기 회귀 경로 생성을 통해 2차원 픽셀 공간에서 탐색 상호작용을 직접 학습하도록 하는 대체 방법을 제안합니다. 주어진 언어 지시와 과거 관찰 내용을 바탕으로, 본 모델은 현재 관찰의 하단 중앙에서 시작하여 일련의 픽셀 좌표를 순차적으로 예측하여 경로를 생성합니다. 기존 연구에서는 픽셀-목표 감독 방식이 이산적인 행동 학습보다 우수하다는 것이 입증되었지만, 본 연구의 실험 결과는 픽셀 공간 경로 감독이 VLN 성능을 크게 향상시킨다는 것을 추가로 확인했습니다. 또한, 본 연구에서 개발한 핵심 모델은 상대적으로 제한된 계산 자원과 훈련 데이터로 최첨단 수준의 성능을 달성하는 것을 입증합니다.

Original Abstract

Benefiting from the powerful priors embedded in large-scale pre-training data and the emerging commonsense reasoning ability, large language models (LLMs) have shown unprecedented generalization capabilities in many research fields. Recently, projecting visual embeddings into the language space via vision-language models (VLMs) to achieve sim-toreal and cross-scene generalization has become a prevailing paradigm in the field of Vision-and-Language Navigation in Continuous Environments (VLN-CE). VLN requires an embodied agent to navigate through unseen environments following natural linguistic instructions. We emphasize that a VLN task can be decomposed into a sequence of sub-tasks, each corresponding to a process of 3D spatial interaction with the environments described by instructions such as "walk to the end of the sofa and turn left." However, such spatial interactions involving moving into the image along the direction of depth sensing are puzzling for VLMs as they were predominantly trained on conversations with RGB images. Rather than incorporating depth or 3D geometric information-which VLMs rarely encounter during pretrainingwe propose an alternative approach: fine-tuning VLMs to learn navigation interactions directly in 2D pixel space through autoregressive trajectory generation. Given a linguistic instruction and historical observations, our model sequentially predicts a series of pixel coordinates, drawing a trajectory from the bottom center of the current observation. While prior work has proved that pixel-goal supervision outperforms learning of discrete actions, our experiments further verify that the supervision of pixel-space trajectory significantly enhances VLN performance. Moreover, we demonstrate that our flagship model achieves state-of-the-art level performance with relatively limited computational resources and training data.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!