2603.22690v1 Mar 24, 2026 cs.CV

WiFi2Cap: Wi-Fi CSI를 이용한 사지 수준의 의미 정렬을 통한 의미 있는 동작 캡션 생성

WiFi2Cap: Semantic Action Captioning from Wi-Fi CSI via Limb-Level Semantic Alignment

Tzu-Ti Wei
Tzu-Ti Wei
Citations: 15
h-index: 2
Yu-Chee Tseng
Yu-Chee Tseng
Citations: 31
h-index: 4
Chuobo Huang
Chuobo Huang
Citations: 1
h-index: 1
Jen-Jee Chen
Jen-Jee Chen
Citations: 2
h-index: 1

실내 환경 감지를 위해서는 인간 활동에 대한 개인 정보 보호적인 의미 이해가 중요하지만, 기존의 Wi-Fi CSI 기반 시스템은 주로 자세 추정 또는 미리 정의된 동작 분류에 초점을 맞추고 있으며, 세밀한 언어 생성을 수행하지 않습니다. CSI를 자연어 설명으로 매핑하는 것은 무선 신호와 언어 사이의 의미 격차, 그리고 좌우 구별과 같은 방향에 민감한 모호성 때문에 어려운 과제입니다. 본 연구에서는 Wi-Fi CSI로부터 직접 동작 캡션을 생성하는 세 단계 프레임워크인 WiFi2Cap을 제안합니다. 비전-언어 기반의 튜처(teacher) 모델은 동기화된 비디오-텍스트 쌍으로부터 전이 가능한 지도 학습을 수행하고, CSI 기반의 학생(student) 모델은 튜처의 시각 공간 및 텍스트 임베딩에 맞춰 정렬됩니다. 방향에 민감한 캡션 생성 성능을 향상시키기 위해, 우리는 좌우 대칭에 의한 동작 및 좌우 구별의 모호성을 줄이는 Mirror-Consistency Loss를 도입했습니다. 마지막으로, 사전 학습된 언어 모델은 CSI 임베딩으로부터 동작 설명을 생성합니다. 또한, 본 연구에서는 Wi-Fi 신호로부터 의미 있는 캡션을 생성하기 위한 동기화된 CSI-RGB-문장 데이터셋인 WiFi2Cap Dataset을 공개합니다. 실험 결과는 WiFi2Cap이 BLEU-4, METEOR, ROUGE-L, CIDEr, 및 SPICE 지표에서 기존 방법보다 일관되게 우수한 성능을 보이며, 효과적인 개인 정보 보호적인 의미 감지를 가능하게 함을 보여줍니다.

Original Abstract

Privacy-preserving semantic understanding of human activities is important for indoor sensing, yet existing Wi-Fi CSI-based systems mainly focus on pose estimation or predefined action classification rather than fine-grained language generation. Mapping CSI to natural-language descriptions remains challenging because of the semantic gap between wireless signals and language and direction-sensitive ambiguities such as left/right limb confusion. We propose WiFi2Cap, a three-stage framework for generating action captions directly from Wi-Fi CSI. A vision-language teacher learns transferable supervision from synchronized video-text pairs, and a CSI student is aligned to the teacher's visual space and text embeddings. To improve direction-sensitive captioning, we introduce a Mirror-Consistency Loss that reduces mirrored-action and left-right ambiguities during cross-modal alignment. A prefix-tuned language model then generates action descriptions from CSI embeddings. We also introduce the WiFi2Cap Dataset, a synchronized CSI-RGB-sentence benchmark for semantic captioning from Wi-Fi signals. Experimental results show that WiFi2Cap consistently outperforms baseline methods on BLEU-4, METEOR, ROUGE-L, CIDEr, and SPICE, demonstrating effective privacy-friendly semantic sensing.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!