CALM-AH: ABAW11 데이터셋 기반의 신뢰도 게이팅 멀티 전문가 합의를 활용한 다중 모달 앙상블 모델을 이용한 비디오 레벨에서의 양면성 및 망설임 인식
CALM-AH: An ABAW11-Calibrated Multimodal Ensemble with Reliability-Gated Multi-Expert Consensus for Video-Level Ambivalence and Hesitancy Recognition
양면성과 망설임(A/H)은 언어, 음성, 표정 활동 및 기타 비언어적 신호를 통해 나타나는 미묘한 행동 상태입니다. ABAW11 A/H 비디오 인식 챌린지는 시스템이 자연스러운 인터뷰 비디오 각각에 대해 이진 A/H 레이블을 할당하도록 요구합니다. 성능은 Macro-F1 점수를 사용하여 측정하며, A/H 샘플과 No-A/H 샘플의 인식이 동등한 중요성을 갖도록 합니다. 본 논문에서는 텍스트, 음향, 시각 및 파생된 행동 통계적 특징을 결합하는 다중 모달 앙상블 모델인 CALM-AH를 제시합니다. 이러한 특징들의 15가지 조합을 구성하고, 각 조합에 대해 검증 데이터셋에서의 이진 교차 엔트로피 손실을 사용하여 세 가지 분류기 중에서 가장 성능이 좋은 것을 선택하고, 검증 Macro-F1 점수를 최적화하기 위해 의사 결정 임계값을 조정합니다. 이렇게 얻어진 이진 결정을 BROTHER 모델에서 전송된 고정된 가중치를 사용하여 결합합니다. 또한, 초기에 예측한 결과를 세 가지 상호 보완적인 수정 전문가(CALM-AH, AffectGPT 및 GPT 기반 의미 검증기)를 통해 결합하는 '신뢰도 게이팅 멀티 전문가 합의'(Reliability-Gated Multi-Expert Consensus, RG-MEC) 기법을 도입합니다. 초기 시스템은 기본 예측값을 제공하며, 이 레이블은 세 가지 수정 전문가가 모두 동일한 대체 클래스를 지지할 때만 변경됩니다. 그렇지 않으면 원래 예측값이 유지됩니다. 이러한 만장일치 기반 설계는 개별 전문가의 오류 영향을 제한하는 동시에, 작업 관련, 다중 모달 감정 및 의미-화용론적 증거가 완전히 일관될 경우 양방향 수정을 허용합니다. 참가자 분리된 ABAW11 데이터셋에서 CALM-AH 모델은 Macro-F1 점수 0.7525를 달성했으며, 전체 RG-MEC 시스템은 0.7771의 Macro-F1 점수를 달성했습니다.
Ambivalence and hesitancy (A/H) are subtle behavioural states that may be expressed through language, voice, facial activity, and other non-verbal cues. The ABAW11 A/H Video Recognition Challenge asks systems to assign a binary A/H label to each naturalistic interview video. Performance is measured using Macro-F1 so that recognition of both A/H and No-A/H samples receives equal importance. We present CALM-AH, a multimodal ensemble that combines textual, acoustic, visual, and derived behavioural-statistical features. We construct 15 non-empty combinations of these feature branches. For each combination, we select the best of three classifier families using validation binary cross-entropy and optimise its decision threshold for validation Macro-F1. The resulting binary decisions are combined using fixed hard-voting weights transferred from BROTHER. We further introduce Reliability-Gated Multi-Expert Consensus(RG-MEC), an anchor-preserving decision-level ensemble that combines an initial prediction with three complementary correction experts: CALM-AH, AffectGPT, and a GPT-based semantic verifier. The initial system provides the default prediction. Its label is overridden only when all three correction experts unanimously support the same alternative class; otherwise, the anchor prediction is retained. This unanimity-gated design limits the influence of isolated expert errors while permitting bidirectional correction when task-specific, multimodal-affective, and semantic-pragmatic evidence are fully consistent. On the participant-disjoint ABAW11 dataset, CALM-AH achieves a Macro-F1 of 0.7525, and the complete RG-MEC system achieves 0.7771.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.