AuricularWorld: 계층적 행동 기반 세계 모델링을 활용한 CT 스캔 이미지의 세밀한 귀 구조 분할
AuricularWorld: Hierarchical Action-Guided World Modeling for Fine-Grained Auricular Structure Segmentation from CT Scans
CT 영상에서 귀 구조를 세밀하게 분할하는 것은 어려운데, 그 이유는 귀가 이미지 영역에서 작은 부분을 차지하고, 연골 경계가 매우 불규칙하며, 연골과 주변 조직 사이의 경계가 종종 모호하기 때문입니다. 또한, 임상 데이터에는 연골을 포함하는 복합 구조와 인접한 피부를 모두 포함하는 라벨이 존재하여 중첩되는 경우가 많습니다. 본 연구에서는 기존의 순방향 예측 방식과는 다른 반복적인 해부학적 추론을 가능하게 하는 세계 모델 기반의 분할 프레임워크를 제안합니다. 이 프레임워크는 인코더-디코더 구조를 기반으로 하며, 중간 잠재 공간에 결정론적인 순환 상태 공간 모델을 도입합니다. 다중 스케일 인코더 특징과 부분적으로 디코딩된 표현을 결합하여 구조적 정보를 생성하고, 이를 통해 잠재 변동을 초기화합니다. 추론 과정에서, 모델은 ground-truth 정보 없이 세 단계의 잠재 변수 전파를 수행합니다. 계층적인 해부학적 행동이 순환 상태를 업데이트하고 잠재 표현을 점진적으로 개선합니다. 결과적으로 얻어진 잠재 궤적은 디코더로 투영되어 고해상도 특징과 결합되어 최종 분할 결과를 생성합니다. 신뢰성 있는 잠재 변수 전환 학습을 위해, 전경의 희소성, 누락된 해부학적 그룹, 그리고 추가 및 제거 연산 간의 불균형 문제를 해결하는 균형 잡힌 계층적 행동 목표 함수를 도입했습니다. 광범위한 실험 결과는 제안된 프레임워크가 CT 이미지에서 작고 불규칙하며 중첩되는 귀 구조에 대해 분할 정확도를 지속적으로 향상시키며, HD95 지표를 43% 이상 감소시킬 수 있음을 보여줍니다. 이러한 결과는 어려운 의료 영상 분할 작업에서 잠재 세계 모델 기반 추론의 효과성을 입증합니다.
Fine-grained segmentation of auricular structures in CT is challenging because the ear occupies a small image region, cartilage boundaries are highly irregular, and interfaces between cartilage and surrounding soft tissues are often ambiguous. Clinical annotations may also include both composite structures containing cartilage and adjacent skin and their corresponding cartilage-only regions, producing nested and overlapping labels. We propose a world-model-based segmentation framework that enables iterative anatomical reasoning beyond conventional feed-forward prediction. Built on an encoder-decoder architecture, the framework introduces a deterministic recurrent state-space model into the intermediate latent space. Multi-scale encoder features and partially decoded representations are fused to form a structural observation that initializes the latent dynamics. During inference, the model performs a three-step latent rollout without ground-truth guidance. Hierarchical anatomical actions update the recurrent state and progressively refine the latent representation. The resulting latent trajectory is projected back into the decoder and combined with high-resolution features to produce the final segmentation. To learn reliable latent transitions, we introduce a balanced hierarchical action objective that addresses foreground sparsity, missing anatomical groups, and imbalance between add and remove operations. Extensive experiments show that the proposed framework consistently improves segmentation accuracy and reduces HD95 by more than 43% for small, irregular, and overlapping auricular structures in CT. These results demonstrate the effectiveness of latent world-model reasoning for challenging medical image segmentation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.