2608.10827v1 Aug 11, 2026 cs.CV

MIRA: 의료 영상 반사를 통한 능동적 진단

MIRA: Medical Image Reflection for Agentic Diagnosis

Xiaozhong Ji
Xiaozhong Ji
Citations: 9
h-index: 2
Qingwen Liu
Qingwen Liu
Citations: 997
h-index: 16
Shengzhi Wang
Shengzhi Wang
Citations: 17
h-index: 2
Jun Yang
Jun Yang
Citations: 28
h-index: 4
Kai Wu
Kai Wu
Citations: 19
h-index: 2
Yiwen Ye
Yiwen Ye
Citations: 456
h-index: 11
Ziyan Chen
Ziyan Chen
Citations: 0
h-index: 0
Mingliang Xiong
Mingliang Xiong
Citations: 494
h-index: 12
Wen Fang
Wen Fang
Citations: 517
h-index: 10
Mingqing Liu
Mingqing Liu
Citations: 597
h-index: 14
Mengyuan Xu
Mengyuan Xu
Citations: 47
h-index: 2
Miaoxuan Shan
Miaoxuan Shan
Citations: 0
h-index: 0
Caiyan Liu
Caiyan Liu
Citations: 0
h-index: 0
Bin He
Bin He
Citations: 8
h-index: 2

의료 시각 에이전트는 이미지를 검사하고 외부 지식을 검색하기 위해 도구를 사용할 수 있지만, 무분별한 도구 사용은 노이즈가 있거나 오해를 불러일으킬 수 있는 증거를 도입할 수 있습니다. 따라서 신뢰할 수 있는 진단은 추가적인 관찰을 얻는 것뿐만 아니라, 도구 사용이 필요한지 여부와 결과적으로 얻어진 증거가 현재 가설을 뒷받침하는지를 확인하는 것을 필요로 합니다. 본 논문에서는 자율적인 증거 탐색 및 반사적 검증을 위한 의료 시각 진단 프레임워크인 MIRA (Medical Image Reflection for Agentic Diagnosis)를 소개합니다. MIRA는 이미지 처리 작업(확대, 영역 지정, 강조 표시, 회전, 측정 등)과 웹 검색을 동적으로 호출하면서, 획득된 증거의 관련성과 일관성을 평가합니다. 우리는 두 단계의 학습 전략을 통해 MIRA를 개발했습니다. 첫째, 도구 기반의 몬테카를로 트리 탐색 데이터 엔진은 다양한 진단 가설을 탐색하고 시각적 영역 지정 정확도와 의미론적 일관성을 동시에 검증하여 지도 학습을 위한 학습 경로를 구축합니다. 둘째, 강화 학습은 온라인 반사 원칙 진화 과정을 통해 의사 결정을 더욱 개선합니다. 실패 사례는 후보 원칙으로 추출되며, 보류된 데이터에 대한 보상 향상을 가져오는 원칙만 유지됩니다. MIRA는 9개의 의료 시각 추론 벤치마크에서 평균 64.73점을 달성하여 Qwen3-VL-8B 기반 모델을 7.44점 개선했습니다. 또한, 유용한 도구 사용 판단률은 56.2%에서 73.8%로 증가하고, 부정적인 판단률은 8.9%에서 1.6%로 감소했습니다. 질적 분석 결과, MIRA는 증거를 재검토하고, 조기에 내린 결론을 수정하며, 도구 사용 전략을 조정할 수 있음을 알 수 있습니다. 프로젝트 페이지: https://MIRA-VL.github.io/

Original Abstract

Medical visual agents can use tools to inspect images and retrieve external knowledge, but indiscriminate tool use may introduce noisy or misleading evidence. Reliable diagnosis therefore requires not only acquiring additional observations, but also verifying whether tool actions are necessary and whether the resulting evidence supports the current hypothesis. We introduce MIRA (Medical Image Reflection for Agentic Diagnosis), a medical visual diagnostic framework for autonomous evidence search and reflective verification. MIRA dynamically invokes image-processing operations, including zooming, grounding, pointing, rotation, and measurement, as well as web search, while evaluating the relevance and consistency of the acquired evidence. We develop MIRA through a two-stage training strategy. First, a tool-augmented Monte Carlo Tree Search data engine explores diverse diagnostic hypotheses and jointly verifies visual grounding accuracy and semantic consistency to construct supervised fine-tuning trajectories. Second, reinforcement learning further improves decision-making through online reflective principle evolution: failure cases are distilled into candidate principles, and only principles that improve held-out rollout rewards are retained. Across nine medical visual reasoning benchmarks, MIRA achieves an average score of 64.73, improving its Qwen3-VL-8B backbone by 7.44 points. It also increases useful tool-use judgments from 56.2% to 73.8% and reduces harmful judgments from 8.9% to 1.6%. Qualitative analyses show that MIRA can re-examine evidence, correct premature conclusions, and adapt its tool-use strategy. Project page: https://MIRA-VL.github.io/

0 Citations
0 Influential
8 Altmetric
40.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!