2607.29310v1 Jul 31, 2026 cs.CV

CALM-AH: ABAW11 데이터셋 기반의 신뢰도 게이팅 멀티 전문가 합의를 활용한 다중 모달 앙상블 모델을 이용한 비디오 레벨에서의 양면성 및 망설임 인식

CALM-AH: An ABAW11-Calibrated Multimodal Ensemble with Reliability-Gated Multi-Expert Consensus for Video-Level Ambivalence and Hesitancy Recognition

Zongyuan Ge
Zongyuan Ge
Citations: 140
h-index: 4
Mingjian Liang
Mingjian Liang
Citations: 300
h-index: 3
Richard Attfield
Richard Attfield
Citations: 2
h-index: 1
Xuelian Cheng
Xuelian Cheng
Citations: 151
h-index: 5
Pamela Carreno-Medrano
Pamela Carreno-Medrano
Citations: 30
h-index: 4
Wenzhuo Sun
Wenzhuo Sun
Citations: 3
h-index: 1

양면성과 망설임(A/H)은 언어, 음성, 표정 활동 및 기타 비언어적 신호를 통해 나타나는 미묘한 행동 상태입니다. ABAW11 A/H 비디오 인식 챌린지는 시스템이 자연스러운 인터뷰 비디오 각각에 대해 이진 A/H 레이블을 할당하도록 요구합니다. 성능은 Macro-F1 점수를 사용하여 측정하며, A/H 샘플과 No-A/H 샘플의 인식이 동등한 중요성을 갖도록 합니다. 본 논문에서는 텍스트, 음향, 시각 및 파생된 행동 통계적 특징을 결합하는 다중 모달 앙상블 모델인 CALM-AH를 제시합니다. 이러한 특징들의 15가지 조합을 구성하고, 각 조합에 대해 검증 데이터셋에서의 이진 교차 엔트로피 손실을 사용하여 세 가지 분류기 중에서 가장 성능이 좋은 것을 선택하고, 검증 Macro-F1 점수를 최적화하기 위해 의사 결정 임계값을 조정합니다. 이렇게 얻어진 이진 결정을 BROTHER 모델에서 전송된 고정된 가중치를 사용하여 결합합니다. 또한, 초기에 예측한 결과를 세 가지 상호 보완적인 수정 전문가(CALM-AH, AffectGPT 및 GPT 기반 의미 검증기)를 통해 결합하는 '신뢰도 게이팅 멀티 전문가 합의'(Reliability-Gated Multi-Expert Consensus, RG-MEC) 기법을 도입합니다. 초기 시스템은 기본 예측값을 제공하며, 이 레이블은 세 가지 수정 전문가가 모두 동일한 대체 클래스를 지지할 때만 변경됩니다. 그렇지 않으면 원래 예측값이 유지됩니다. 이러한 만장일치 기반 설계는 개별 전문가의 오류 영향을 제한하는 동시에, 작업 관련, 다중 모달 감정 및 의미-화용론적 증거가 완전히 일관될 경우 양방향 수정을 허용합니다. 참가자 분리된 ABAW11 데이터셋에서 CALM-AH 모델은 Macro-F1 점수 0.7525를 달성했으며, 전체 RG-MEC 시스템은 0.7771의 Macro-F1 점수를 달성했습니다.

Original Abstract

Ambivalence and hesitancy (A/H) are subtle behavioural states that may be expressed through language, voice, facial activity, and other non-verbal cues. The ABAW11 A/H Video Recognition Challenge asks systems to assign a binary A/H label to each naturalistic interview video. Performance is measured using Macro-F1 so that recognition of both A/H and No-A/H samples receives equal importance. We present CALM-AH, a multimodal ensemble that combines textual, acoustic, visual, and derived behavioural-statistical features. We construct 15 non-empty combinations of these feature branches. For each combination, we select the best of three classifier families using validation binary cross-entropy and optimise its decision threshold for validation Macro-F1. The resulting binary decisions are combined using fixed hard-voting weights transferred from BROTHER. We further introduce Reliability-Gated Multi-Expert Consensus(RG-MEC), an anchor-preserving decision-level ensemble that combines an initial prediction with three complementary correction experts: CALM-AH, AffectGPT, and a GPT-based semantic verifier. The initial system provides the default prediction. Its label is overridden only when all three correction experts unanimously support the same alternative class; otherwise, the anchor prediction is retained. This unanimity-gated design limits the influence of isolated expert errors while permitting bidirectional correction when task-specific, multimodal-affective, and semantic-pragmatic evidence are fully consistent. On the participant-disjoint ABAW11 dataset, CALM-AH achieves a Macro-F1 of 0.7525, and the complete RG-MEC system achieves 0.7771.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!