2604.10541v3 Apr 12, 2026 cs.CV

이질적인 데이터 세트를 이용한 구조화된 의미 매핑을 통한 얼굴 표정 단위 및 표현의 양방향 학습

Bidirectional Learning of Facial Action Units and Expressions via Structured Semantic Mapping across Heterogeneous Datasets

Yong Li
Yong Li
Citations: 465
h-index: 4
Shiguang Shan
Shiguang Shan
Citations: 27
h-index: 2
Jia Li
Jia Li
Citations: 63
h-index: 5
Yu Zhang
Yu Zhang
Citations: 18
h-index: 2
Yin Chen
Yin Chen
Citations: 129
h-index: 4
Zhenzhen Hu
Zhenzhen Hu
Citations: 45
h-index: 4
Richang Hong
Richang Hong
Citations: 49
h-index: 4
Meng Wang
Meng Wang
Citations: 17
h-index: 3

얼굴 표정 단위(AU) 감지 및 얼굴 표정(FE) 인식은 미세한 근육 활성화와 거친 수준의 전반적인 정서 상태를 나타내는 감정적 얼굴 행동 작업으로 함께 볼 수 있습니다. 이러한 두 가지는 내재적으로 의미론적 상관 관계가 있지만, 기존 연구에서는 주로 AU에서 FE로의 지식 전달에 초점을 맞추고 있으며, 양방향 학습은 충분히 탐구되지 않았습니다. 더욱이, 실제로는 데이터 조건의 이질성으로 인해 AU 및 FE 데이터 세트 간의 주석 방식(프레임 레벨 vs.\ 클립 레벨), 레이블 세분화 수준, 그리고 데이터 가용성과 다양성이 달라 효과적인 공동 학습을 어렵게 만듭니다. 이러한 문제를 해결하기 위해, 우리는 다양한 데이터 도메인과 이질적인 감독 환경에서 AU-FE 양방향 학습을 위한 구조화된 의미 매핑(SSM) 프레임워크를 제안합니다. SSM은 세 가지 핵심 구성 요소로 이루어져 있습니다: (1) 동적 AU 및 FE 비디오에서 통일된 얼굴 표현을 학습하는 공유 시각적 백본; (2) 텍스트 기반 의미 프로토타입(TSP) 모듈을 통한 의미 중재, 이는 학습 가능한 컨텍스트 프롬프트를 사용하여 고정된 텍스트 설명을 기반으로 구조화된 의미 프로토타입을 구축하고, 공유 의미 공간에서 감독 및 교차 작업 정렬을 수행합니다; (3) FACS에서 파생된 사전 지식을 통합하고, 명시적인 지식 전달을 위해 텍스트 의미 공간에서 데이터 적응형 양방향 연관 행렬을 학습하는 동적 사전 매핑(DPM) 모듈. 인기 있는 AU 감지 및 FE 인식 벤치마크에 대한 광범위한 실험 결과는 SSM이 단일 작업 및 다중 작업 기준 모델보다 일관되게 뛰어난 성능을 보이며, 특정 작업에 최적화된 방법과 경쟁력 있는 성능을 달성함을 보여줍니다. 또한 FE에서 AU로의 학습 결과는 전반적인 표정 의미론이 이질적인 데이터 세트 전체에서 미세한 AU 학습에 유용한 감독 신호를 제공한다는 것을 보여줍니다.

Original Abstract

Facial action unit (AU) detection and facial expression (FE) recognition can be jointly viewed as affective facial behavior tasks, representing fine-grained muscular activations and coarse-grained holistic affective states, respectively. Despite their inherent semantic correlation, existing studies predominantly focus on knowledge transfer from AUs to FEs, while bidirectional learning remains insufficiently explored. In practice, this challenge is further compounded by heterogeneous data conditions, where AU and FE datasets differ in annotation paradigms (frame-level vs.\ clip-level), label granularity, and data availability and diversity, hindering effective joint learning. To address these issues, we propose a Structured Semantic Mapping (SSM) framework for bidirectional AU--FE learning under different data domains and heterogeneous supervision. SSM consists of three key components: (1) a shared visual backbone that learns unified facial representations from dynamic AU and FE videos; (2) semantic mediation via a Textual Semantic Prototype (TSP) module, which constructs structured semantic prototypes from fixed textual descriptions with learnable context prompts for supervision and cross-task alignment in a shared semantic space; and (3) a Dynamic Prior Mapping (DPM) module that incorporates FACS-derived prior knowledge and learns data-adaptive bidirectional association matrices in the textual semantic space for explicit knowledge transfer. Extensive experiments on popular AU detection and FE recognition benchmarks show that SSM consistently outperforms its single-task and multi-task baselines and achieves competitive performance against task-specific methods. The FE-to-AU results further show that holistic expression semantics provides useful supervision for fine-grained AU learning across heterogeneous datasets.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!