지식 기반 분리 학습: 원자적 동작을 활용한 행동 인식
Knowledge-guided Disentanglement with Atomic Actions for Action Recognition
복잡한 장면에서의 행동 인식은 종종 여러 개의 동시에 발생하는 미세한 행동들을 포함하며, 이는 내부적인 행동 구조를 모델링하는 데 어려움을 야기합니다. 대부분의 기존 방법들은 전체적인 표현에 의존하는데, 이는 미묘한 상호작용과 세밀한 의미를 포착하기에는 부족합니다. 최근 프롬프트 기반 접근 방식은 분리 학습을 도입했지만, 명시적인 의미 지침이 부족하며, 시각적 또는 구조적 단서만을 사용하는 방법들은 여전히 粗粒度입니다. 본 논문에서는 원자적 동작을 활용한 지식 기반 분리 학습(Knowledge-guided Disentanglement with Atomic Actions, KDA) 방법을 제안합니다. 이는 세밀한 의미 지식을 활용하여 행동 표현을 향상시키고 더욱 정확한 분리 학습을 가능하게 합니다. 구체적으로, 우리는 대규모 언어 모델(LLM)을 사용하여 동작 레이블을 원자적 동작으로 분해하고, 명시적인 시공간적 의미를 제공합니다. 지식 주입 모듈(Knowledge Injection Module, KIM)은 먼저 원자적 동작 지식을 비디오 특징에 통합합니다. 이 향상된 표현을 바탕으로, 지식 분리 모듈(Knowledge Disentanglement Module, KDM)은 원자적 동작 지식을 더욱 세밀하게 분리하여 행동 분리 학습에 대한 더 정확한 의미 지침을 제공합니다. 또한, KDM 내에서 지식 구성 요소를 명확하게 분리하도록 유도하는 지식 분리 손실(Knowledge Disentanglement Loss, KD Loss)을 도입했습니다. 광범위한 실험 결과는 KDA가 특징의 구별력을 향상시키고 다중 레이블 행동 인식 벤치마크에서 최첨단 성능을 달성한다는 것을 보여줍니다. 또한, KIM과 KDM은 다른 방법에도 쉽게 통합될 수 있으며, 이는 높은 일반성을 입증합니다.
Action recognition in complex scenes often involves multiple concurrent fine-grained actions, making it challenging to model internal action structures. Most existing methods rely on holistic representations, which are insufficient for capturing subtle interactions and fine-grained semantics. While recent prompt-based approaches introduce disentanglement, they lack explicit semantic guidance, and methods based solely on visual or structured cues remain coarse-grained. In this paper, we propose Knowledge-guided Disentanglement with Atomic Actions (KDA), which leverages fine-grained semantic knowledge to enhance action representations and enable more precise disentanglement. Specifically, we use Large Language Models (LLMs) to decompose action labels into atomic actions, providing explicit spatial-temporal semantics. A Knowledge Injection Module (KIM) first integrates atomic action knowledge into video features. Based on this enhanced representation, a Knowledge Disentanglement Module (KDM) further disentangles atomic action knowledge to produce more precise semantic guidance for action disentanglement. A Knowledge Disentanglement Loss (KD Loss) is introduced to encourage clearer disentanglement of knowledge components within KDM. Extensive experiments demonstrate that KDA improves feature discriminability and achieves state-of-the-art performance on multi-label action recognition benchmarks. Moreover, KIM and KDM can be readily integrated into other methods, demonstrating strong generality.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.