DiffImaginE: 확산 모델을 활용한 개체 유형 검증
DiffImaginE: Imagine to Verify Entity Types with Diffusio
다중 모드 명명 개체 인식(MNER)은 각 후보 구간과 개체 유형 가설이 텍스트 및 시각적 증거에 의해 뒷받침되는지 여부를 결정합니다. 기존의 이미지 기반 검증 방식은 각 (구간, 유형) 쌍을 하나의 예측된 시각적 특징으로 매핑하여 다양한 시각적 표현을 단일 프로토타입으로 압축하고 명시적인 확률적 의미론 없이 호환성 점수를 제공합니다. 본 논문에서는 MNER 유형 검증을 조건부 잠재 확산 추론 문제로 정의하는 DiffImaginE를 소개합니다. 구간 위치 정보를 포함하는 시각적 증거가 주어지면, 유형에 따라 조정된 디노이징 모델은 표준화된 잠재 공간에 주입된 노이즈를 예측합니다. 결과적으로 생성되는 디노이징 오차는 유형 조건부 음의 로그 가능도를 추정하며, 이를 통해 경쟁하는 유형 가설을 관찰 데이터에 대한 설명력으로 순위를 매길 수 있습니다. DiffImaginE는 표준적인 다중 모드 인코더 구조를 유지하고, 결정론적 검증기를 Min-SNR 가중치를 사용하여 학습된 분류기 없는 안내 확산 점수기로 대체합니다. 각 유형별 확산 점수를 분류 로그 확률로 직접 감독하고, 노이즈 레벨 간의 집계를 학습하며, 반대 샘플링을 사용하여 몬테카를로 비교의 분산을 줄입니다. 분석 결과, 분류기 없는 안내가 유도된 유형 사후 분포를 더욱 명확하게 만들고, 특정 조건에서 반대 페어링이 디노이징 모델 비용에 미치는 영향과 분산 감소 효과를 설명합니다. Twitter-2015 및 Twitter-2017 데이터셋에서의 실험 결과는 동일한 인코더, 보조 목표, 평가 프로토콜을 사용한 결정론적 ImaginE 모델 대비 일관된 성능 향상을 보여주며, 이는 추가적인 분석과 통계적 유의성 검증을 통해 뒷받침됩니다.
Multimodal named entity recognition (MNER) determines whether each candidate span and entity-type hypothesis is supported by joint textual and visual evidence. Existing imagine-and-compare verifiers map each (span, type) pair to one predicted visual feature, compressing diverse visual realisations into a single prototype and providing a compatibility score without explicit probabilistic semantics. We introduce DiffImaginE, which formulates MNER type verification as conditional latent diffusion inference. Given span-localised visual evidence, a type-conditioned denoiser predicts noise injected into its standardised latent. The resulting denoising error provides an ELBO-consistent surrogate for type-conditional negative log-likelihood, allowing competing type hypotheses to be ranked by how well they explain the observation. DiffImaginE retains a standard multimodal encoder stack and replaces the deterministic verifier with a classifier-free-guided diffusion scorer trained using Min-SNR weighting. We directly supervise per-type diffusion scores as classification logits, learn aggregation across noise levels, and use antithetic sampling to reduce Monte Carlo comparison variance. Our analysis shows that classifier-free guidance sharpens the induced type posterior and characterises when antithetic pairing reduces variance at equal denoiser cost. Experiments on Twitter-2015 and Twitter-2017 show consistent gains over a matched deterministic ImaginE control under the same encoder, auxiliary objectives, and evaluation protocol, supported by ablations and paired significance tests.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.