2601.06550v1 Jan 10, 2026 cs.CV

LLMTrack: 멀티모달 대규모 언어 모델을 활용한 의미 기반 다중 객체 추적

LLMTrack: Semantic Multi-Object Tracking with Multi-modal Large Language Models

Pan Liao
Pan Liao
Citations: 45
h-index: 4
Feng Yang
Feng Yang
Citations: 197
h-index: 4
Di Wu
Di Wu
Citations: 41
h-index: 4
Jinwen Yu
Jinwen Yu
Citations: 13
h-index: 2
Dingwen Zhang
Dingwen Zhang
Citations: 11
h-index: 2
Wang Zhao
Wang Zhao
Citations: 22
h-index: 3

기존의 다중 객체 추적 (MOT) 시스템은 위치 추적 및 객체 연결에서 뛰어난 정확도를 보여주며, 객체가 '어디에' 있고 '누구'인지 효과적으로 파악합니다. 그러나 이러한 시스템은 종종 자폐적인 관찰자처럼 작동하며, 기하학적 경로를 추적하는 데는 능숙하지만 객체의 행동 뒤에 숨겨진 의미적인 '무엇'과 '왜'에 대해서는 이해하지 못합니다. 본 논문에서는 기하학적 인식과 인지적 추론 간의 간극을 해소하기 위해, 의미 기반 다중 객체 추적 (SMOT)을 위한 새로운 통합 프레임워크인 **LLMTrack**을 제안합니다. 우리는 강력한 위치 추적 능력과 심층적인 이해 능력을 분리하는 생체 모방 설계 철학을 채택하여, Grounding DINO를 '눈'으로, LLaVA-OneVision 멀티모달 대규모 언어 모델을 '뇌'로 사용합니다. 우리는 객체 수준의 상호 작용 특징과 비디오 수준의 문맥 정보를 통합하는 시공간 융합 모듈을 도입하여, 대규모 언어 모델 (LLM)이 복잡한 궤적을 이해할 수 있도록 합니다. 또한, 시각적 정렬, 시간적 미세 조정, 그리고 LoRA를 통한 의미 주입을 포함하는 점진적인 세 단계 훈련 전략을 설계하여, 방대한 모델을 추적 도메인에 효율적으로 적용합니다. BenSMOT 벤치마크에서 수행한 광범위한 실험 결과, LLMTrack은 최첨단 성능을 달성했으며, 객체 설명, 상호 작용 인식, 비디오 요약에서 기존 방법보다 훨씬 뛰어난 성능을 보이며, 동시에 견고한 추적 안정성을 유지합니다.

Original Abstract

Traditional Multi-Object Tracking (MOT) systems have achieved remarkable precision in localization and association, effectively answering \textit{where} and \textit{who}. However, they often function as autistic observers, capable of tracing geometric paths but blind to the semantic \textit{what} and \textit{why} behind object behaviors. To bridge the gap between geometric perception and cognitive reasoning, we propose \textbf{LLMTrack}, a novel end-to-end framework for Semantic Multi-Object Tracking (SMOT). We adopt a bionic design philosophy that decouples strong localization from deep understanding, utilizing Grounding DINO as the eyes and the LLaVA-OneVision multimodal large model as the brain. We introduce a Spatio-Temporal Fusion Module that aggregates instance-level interaction features and video-level contexts, enabling the Large Language Model (LLM) to comprehend complex trajectories. Furthermore, we design a progressive three-stage training strategy, Visual Alignment, Temporal Fine-tuning, and Semantic Injection via LoRA to efficiently adapt the massive model to the tracking domain. Extensive experiments on the BenSMOT benchmark demonstrate that LLMTrack achieves state-of-the-art performance, significantly outperforming existing methods in instance description, interaction recognition, and video summarization while maintaining robust tracking stability.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!