TAVR-VLM: 위험 조건 기반 인과적 연결을 통한 환각 방지 보고서 생성
TAVR-VLM: Risk-Conditioned Causal Grounding for Hallucination-Resistant Report Generation
경고막혈관 판막 치환술(TAVR) 계획은 세심한 다중 모드 추론이 필요합니다. 그러나, 다중 모드 대규모 언어 모델(MLLM)을 이 고위험 영역에 적용하는 데는 진단적 환각 문제가 심각하게 걸림돌이 됩니다. 이러한 문제를 해결하기 위해, TAVR-VLM이라는 새로운 프레임워크가 제안됩니다. 이는 위험 조건 기반 인과적 연결 주의 메커니즘(R-CGA)을 특징으로 하며, 모델 내부적으로 ``위험 → 영역 → 단어'' 구조적 연결 경로를 구현합니다. R-CGA는 다중 모드 입력을 인과적 위험 병목 현상으로 압축하여, 밀집된 시각적 특징을 전역적인 위험 마스크로 정제합니다. 자기 회귀 생성 과정에서, 지원 벡터 투영 기반의 인과적 일관성 목표 함수는 위험이 정의된 지원 마스크 내에서 토큰 수준의 연결성을 제약합니다. 1,482명의 환자 코호트인 $ ext{M}^3 ext{TAVR}$ 데이터셋에 대한 평가 결과, TAVR-VLM은 새로운 최고 성능을 달성했습니다. AUROC는 0.896으로 상승했으며, CIDEr 점수는 0.936으로 향상되었고, 환각 발생률은 8.1%로 크게 감소하여 증거 기반 수술 인공지능의 해석 가능성을 향상시켰습니다.
Transcatheter Aortic Valve Replacement (TAVR) planning requires meticulous multimodal reasoning. However, adapting Multimodal Large Language Models (MLLMs) to this high-stakes domain is severely impeded by diagnostic hallucinations, where generated text lacks anatomical grounding. To address this, TAVR-VLM is introduced: a novel framework featuring Risk-Conditioned Causal Grounding Attention (R-CGA) that instantiates a model-internal ``Risk $\rightarrow$ Region $\rightarrow$ Word'' structural grounding pathway. R-CGA compresses multimodal inputs into a causal risk bottleneck, purifying dense visual features into a global risk mask. During autoregressive generation, a support-projected causal consistency objective constrains token-level grounding within the risk-defined support mask. Evaluated on $\text{M}^3\text{TAVR}$, a comprehensive 1,482-patient cohort, TAVR-VLM establishes a new state-of-the-art. It achieves an AUROC of 0.896, boosts CIDEr to 0.936, and drastically reduces the hallucination rate to 8.1\%, thereby improving interpretability for evidence-based surgical AI.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.