훈련 없이 신경 방사선 이미지 분석을 위한 능동적인 대규모 언어 모델
Agentic Large Language Models for Training-Free Neuro-Radiological Image Analysis
최첨단 대규모 언어 모델(LLM)은 일반적인 시각 질의 응답에서 높은 성능을 보입니다. 그러나 근본적인 한계가 여전히 존재합니다. 현재 아키텍처는 CT 또는 MRI와 같은 입체 의료 이미지를 직접 분석하는 데 필요한 기본적인 3차원 공간 추론 능력이 부족합니다. 새롭게 떠오르는 능동형 인공지능은 이러한 문제에 대한 새로운 해결책을 제시하며, LLM이 내재적인 3차원 처리가 필요 없도록 특수 외부 도구를 활용하고 조율할 수 있도록 합니다. 그러나 복잡하고 다단계 방사선 작업 흐름에서 이러한 능동형 프레임워크의 실현 가능성은 아직 충분히 연구되지 않았습니다. 본 연구에서는 훈련 없이 자동화된 뇌 MRI 분석을 위한 능동형 파이프라인을 제시합니다. GPT-5.1, Gemini 3 Pro, Claude Sonnet 4.5 등 다양한 LLM과 상용 전문 도구를 사용하여 당사의 방법론을 검증한 결과, 시스템은 전처리(두개골 제거, 정렬), 병변 분할(교모종, 뇌막종, 전이), 부피 분석 등 복잡한 엔드투엔드 작업 흐름을 자율적으로 실행합니다. 당사는 단일 스캔 분할 및 부피 보고에서부터 다점 시기 비교를 요구하는 장기 반응 평가에 이르기까지 점진적으로 복잡해지는 방사선학적 작업을 통해 당사의 프레임워크를 평가했습니다. 아키텍처 설계의 영향을 분석하기 위해 단일 에이전트 모델과 여러 에이전트 간의 "도메인 전문가" 협업을 비교했습니다. 마지막으로, 향후 능동형 시스템의 엄격한 평가를 지원하기 위해 공개 BraTS 데이터에서 파생된 이미지-프롬프트-응답 튜플로 구성된 벤치마크 데이터 세트를 소개하고 공개합니다. 당사의 결과는 능동형 인공지능이 훈련이나 미세 조정 없이 도구 사용을 통해 고도의 신경 방사선 이미지 분석 작업을 해결할 수 있음을 보여줍니다.
State-of-the-art large language models (LLMs) show high performance in general visual question answering. However, a fundamental limitation remains: current architectures lack the native 3D spatial reasoning required for direct analysis of volumetric medical imaging, such as CT or MRI. Emerging agentic AI offers a new solution, eliminating the need for intrinsic 3D processing by enabling LLMs to orchestrate and leverage specialized external tools. Yet, the feasibility of such agentic frameworks in complex, multi-step radiological workflows remains underexplored. In this work, we present a training-free agentic pipeline for automated brain MRI analysis. Validating our methodology on several LLMs (GPT-5.1, Gemini 3 Pro, Claude Sonnet 4.5) with off-the-shelf domain-specific tools, our system autonomously executes complex end-to-end workflows, including preprocessing (skull stripping, registration), pathology segmentation (glioma, meningioma, metastases), and volumetric analysis. We evaluate our framework across increasingly complex radiological tasks, from single-scan segmentation and volumetric reporting to longitudinal response assessment requiring multi-timepoint comparisons. We analyze the impact of architectural design by comparing single-agent models against multi-agent "domain-expert" collaborations. Finally, to support rigorous evaluation of future agentic systems, we introduce and release a benchmark dataset of image-prompt-answer tuples derived from public BraTS data. Our results demonstrate that agentic AI can solve highly neuro-radiological image analysis tasks through tool use without the need for training or fine-tuning.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.