2608.14243v1 Aug 14, 2026 cs.CV

제로샷 기반 골격 데이터 활용 액션 예측

Zero-Shot Skeleton-Based Action Anticipation

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Peng-Cheng Yan
Peng-Cheng Yan
Citations: 10
h-index: 1
Qiuxia Lai
Qiuxia Lai
Citations: 455
h-index: 3

액션 예측(Action Anticipation, AA)은 부분적인 관찰만으로 인간 또는 휴머노이드의 현재 진행 중인 행동을 인식하여 로봇이 행동 완료 전에 의도를 예측할 수 있도록 하는 기술입니다. 골격 데이터를 이용한 액션 예측은 효율성 측면에서 장점을 제공하지만, 기존 연구들은 훈련 과정에서 모든 액션 클래스가 존재한다고 가정하기 때문에 실제 환경에서 새롭게 발생하는 액션에 대한 적용 가능성이 제한됩니다. 이러한 문제점을 해결하기 위해 본 연구에서는 새로운 과제인 제로샷 기반 골격 데이터 활용 액션 예측(Zero-Shot Skeleton-Based Action Anticipation, ZS-SkAA)을 다룹니다. 이 과제는 제한된 초기 단계의 골격 데이터만을 사용하여 학습되지 않은 액션 클래스를 인식해야 하며, 이는 부분적인 관찰, 시간적 동역학, 그리고 제로샷 일반화라는 어려움을 동시에 포함합니다. ZS-SkAA에 대한 기초 연구를 확립하기 위해 다음과 같은 내용을 제시합니다: (1) 공간-시간 특징 추출기와 상호 정보 추정 및 최대화 모듈로 구성된 기본 모델을 개발했습니다. 이 기본 모델은 시각적 특징과 의미론적 클래스 임베딩 간의 상호 정보를 추정하고 최대화하여 여러 modality에 걸쳐 특징을 정렬함으로써 학습되지 않은 클래스로의 일반화를 향상시킵니다. (2) NTU RGB+D 데이터셋을 활용한 벤치마크 프로토콜을 개발하여 ZS-SkAA 평가를 위한 엄격한 기준을 제시합니다. 실험 결과는 제안하는 모델이 ZS-SkAA에 대한 강력한 기본 모델로서의 성능을 입증하며, NTU RGB+D 데이터셋에서 높은 제로샷 정확도를 달성했습니다. 본 연구는 새로운 액션에 대한 일반화가 필요한 실제 시스템을 위한 중요한 연구 방향으로서 ZS-SkAA의 중요성을 강조합니다.

Original Abstract

Action anticipation (AA) aims to recognize ongoing human or humanoids actions from partial observations, enabling robots to predict intentions before the actions are completed. Although skeleton-based AA offers efficiency advantages, existing approaches assume that all action classes are seen during training, which limits their deployment in real-world scenarios where novel actions inevitably arise. To address this gap, we study the new task of Zero-Shot Skeleton-Based Action Anticipation (ZS-SkAA). This task requires recognizing unseen action classes using only limited early-stage skeleton sequences, combining the challenges of partial observations, temporal dynamics, and zero-shot generalization. To establish foundational research for ZS-SkAA, we introduce:(1) A baseline model comprising a spatio-temporal feature extractor and a mutual information estimation and maximization module. This baseline model explicitly aligns partial visual features with semantic class embeddings across modalities by estimating and maximizing their mutual information, enhancing generalization to unseen classes.(2) A benchmark protocol using the NTU RGB+D dataset, which is adapted for rigorous ZS-SkAA evaluation. Experiments demonstrate the effectiveness of our model as a strong baseline for ZS-SkAA, achieving high zero-shot accuracy on NTU RGB+D. This work establishes ZS-SkAA as a vital research direction for real-world systems requiring generalization to novel actions.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!