2603.14276v1 Mar 15, 2026 cs.CV

Tucker 적응 기반의 전일정 다중 환경 지속적인 시각-언어 탐색

All-day Multi-scenes Lifelong Vision-and-Language Navigation with Tucker Adaptation

Xudong Wang
Xudong Wang
Citations: 46
h-index: 3
Lianqing Liu
Lianqing Liu
Citations: 12
h-index: 3
Zhi Han
Zhi Han
Citations: 13
h-index: 3
Zhiyu Liu
Zhiyu Liu
Citations: 58
h-index: 4
Gang Li
Gang Li
Citations: 3,288
h-index: 3
Yao Wang
Yao Wang
Citations: 28
h-index: 3

시각-언어 탐색(VLN) 에이전트를 활용하려면 다양한 환경과 장면에서의 적응이 필수적이지만, 특정 시나리오에 대한 미세 조정은 다른 시나리오에서 심각한 망각 현상을 야기하여 장기적인 활용성을 제한합니다. 본 연구에서는 이러한 문제를 전일정 다중 환경 지속적인 VLN(AML-VLN) 문제로 공식화합니다. 기존의 파라미터 효율적인 어댑터(예: LoRA 및 변형)는 2차원 행렬 형태로 표현되어 여러 장면과 환경에 걸쳐 존재하는 다중 계층의 탐색 지식을 제대로 반영하지 못합니다. 이를 해결하기 위해, 본 연구에서는 다중 계층의 탐색 지식을 고차원 텐서로 표현하고, Tucker 분해를 활용하여 지식을 공유된 부분 공간과 시나리오별 전문가로 분리하는 Tucker Adaptation (TuKA)을 제안합니다. 또한, 공유된 부분 공간을 강화하면서 특정 전문가를 제한하는 분리된 지식 점진적 학습 전략을 도입하여 지속적인 학습을 가능하게 합니다. TuKA를 기반으로, AlldayWalker라는 VLN 에이전트를 개발하여 여러 탐색 시나리오에서 지속적으로 학습하고, 전일정 다중 환경 탐색을 달성합니다. 광범위한 실험 결과, AlldayWalker는 최첨단 모델보다 우수한 성능을 지속적으로 보여줍니다.

Original Abstract

Deploying vision-and-language navigation (VLN) agents requires adaptation across diverse scenes and environments, but fine-tuning on a specific scenario often causes catastrophic forgetting in others, which severely limits flexible long-term deployment. We formalize this challenge as the all-day multi-scenes lifelong VLN (AML-VLN) problem. Existing parameter-efficient adapters (e.g., LoRA and its variants) are limited by their two-dimensional matrix form, which fails to capture the multi-hierarchical navigation knowledge spanning multiple scenes and environments. To address this, we propose Tucker Adaptation (TuKA), which represents the multi-hierarchical navigation knowledge as a high-order tensor and leverages Tucker decomposition to decouple the knowledge into shared subspaces and scenario-specific experts. We further introduce a decoupled knowledge incremental learning strategy to consolidate shared subspaces while constraining specific experts for decoupled lifelong learning. Building on TuKA, we also develop a VLN agent named AlldayWalker, which continually learns across multiple navigation scenarios, achieving all-day multi-scenes navigation. Extensive experiments show that AlldayWalker consistently outperforms state-of-the-art baselines.

3 Citations
0 Influential
2 Altmetric
13.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!