증거 기반 다중 에이전트 추론을 통한 신뢰성 있는 감정 이미지 설명: 접근 방법
Towards Faithful Sentimental Image Captioning via Evidence-Aware Multi-Agent Reasoning
감정 이미지 설명(SIC)은 감정 표현과 시각적 충실도 간의 균형을 요구합니다. 기존 방법들은 종종 이러한 상충 관계 때문에 어려움을 겪으며, 이는 충분한 지역 기반 부족 및 감정 검증 메커니즘의 부재로 인한 환각 현상을 야기할 수 있습니다. 이러한 한계를 극복하기 위해, 우리는 신뢰성 있고 증거 기반의 감정 이미지 설명을 위한 감정-증거 인식 다중 에이전트 시스템인 SEA-Cap을 제안합니다. SEA-Cap은 감정 증거 채굴기를 통합하여 구조화된 지역적 정서 단서를 추출하고, 감정 제어를 전역 속성에서 검증 가능한 객체 수준의 증거로 이동시킵니다. 이 증거를 활용하여, 우리의 프레임워크는 생성기, 환각 검사기 및 중재기가 공유 블랙보드를 통해 반복적으로 설명을 개선하는 협력적 워크플로우를 구축합니다. SEA-Cap은 생성된 콘텐츠를 채굴된 시각적 증거와 명시적으로 비교함으로써 감정 정확성과 사실 일관성을 모두 보장합니다. 두 개의 벤치마크 데이터 세트에 대한 광범위한 실험 결과, SEA-Cap이 환각 현상을 효과적으로 줄이고 최첨단 성능을 달성한다는 것을 보여줍니다.
Sentimental Image Captioning (SIC) requires balancing emotional expression with visual fidelity. Existing methods often struggle with this trade-off, leading to hallucinations due to insufficient local grounding and the lack of sentimental verification mechanisms. To address these limitations, we propose SEA-Cap, a Sentiment-Evidence-Aware Multi-Agent System for faithful and evidence-grounded sentimental image captioning. SEA-Cap incorporates a Sentiment Evidence Miner that extracts structured, local affective cues to shift sentiment control from global attributes to verifiable object-level evidence. Leveraging this evidence, our framework orchestrates a collaborative workflow where a Generator, Hallucination Checker, and Arbitrator iteratively refine captions via a shared blackboard. By explicitly auditing generated content against mined visual evidence, SEA-Cap ensures both sentiment accuracy and factual consistency. Extensive experiments on two benchmark datasets demonstrate that SEA-Cap effectively mitigates hallucinations and achieves state-of-the-art performance.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.