CrossScope: 역할 비대칭 세계 모델을 활용한 복합 시야 수술 영상 예측
CrossScope: A Role-Asymmetric World Model for Joint Dual-Scope Surgical Video Prediction
시각적 세계 모델은 일반적으로 단일 관측 데이터 스트림으로부터 미래의 동역학을 학습하지만, 이는 여러 독립적으로 움직이는 관찰자가 존재하는 협력 시스템을 모델링하는 능력을 제한합니다. 본 연구에서는 모-자 내시경 역행성 담췌관 조영술(ERCP)에서 발생하는 이러한 문제를 조사합니다. 이 절차에서는 두 개의 유연한 내시경이 상호 보완적이지만 역할에 따라 다른 시야를 제공하며, 정확하게 교정된 입체 관계를 가지고 있지 않습니다. 기존의 다중 뷰 융합 방식은 대칭적인 정보 교환을 가정하는 반면, 우리는 **역할 비대칭 복합 시야 미래 예측** 문제를 정의합니다. 여기서 각 시점에서 얻은 정보는 예측 대상과 그에 따른 공간적 요구 사항에 따라 선택적으로 이전됩니다. 본 연구에서는 **CrossScope**라는 이중 스트림 수술 세계 모델을 제안합니다. CrossScope는 각 시점의 특성을 유지하면서도, 기하학 정보를 기반으로 하는 잔차 상호 작용을 통해 예측 대상에 특화된 정보 전달을 가능하게 합니다. CrossScope는 두 가지 보완적인 통신 방향을 학습합니다. 모(Mother) 뷰에서 얻은 기하학적 운동 정보는 자(Child) 뷰의 미래 동역학을 안내하며, 반면, 자세가 일치하는 자 뷰의 특징 정보는 유효한 공간적 대응 관계가 확립되었을 때에만 모 뷰 예측을 지원합니다. 이러한 설계 덕분에 각 내시경은 특정 작업과 관련된 정보를 제공하면서도 고유한 시점 표현을 유지할 수 있습니다. 본 연구에서는 이 문제를 평가하기 위해, 동기화된 가상 및 실제 ERCP 영상을 포함하는 복합 시야 벤치마크를 구축했습니다. 시각적 충실도, 구조 보존, 대상 위치 파악, 운동 일관성을 평가하여 성능을 측정했습니다. 실험 결과는 CrossScope가 기존의 강력한 수술 영상 생성 모델보다 우수한 성능을 보임을 입증하며, 다중 관찰자 기반의 시각적 세계 모델링에서 역할 인지적인 정보 전달의 중요성을 강조합니다.
Visual world models typically learn future dynamics from a single observation stream, limiting their ability to model cooperative systems with multiple independently moving observers. We investigate this challenge in Mother--Child endoscopic retrograde cholangiopancreatography (ERCP), where two flexible scopes provide complementary yet role-dependent views without a calibrated stereo relationship. Unlike conventional multi-view fusion that assumes symmetric information exchange, we formulate \textbf{role-asymmetric dual-scope future prediction}, where cross-view evidence is selectively transferred according to the prediction target and its underlying spatial requirements. We propose \textbf{CrossScope}, a dual-stream surgical world model that preserves view-specific experts while enabling target-specific evidence routing through geometry-guided residual interactions. CrossScope learns two complementary communication directions: geometric motion cues from the Mother view guide Child-view future dynamics, while pose-aligned Child appearance supports Mother-view prediction only when valid spatial correspondence is established. This design allows each scope to contribute task-relevant evidence without compromising its view-specific representation. To evaluate this problem, we establish a paired dual-scope benchmark comprising synchronized phantom and real-world ERCP episodes, with evaluations assessing visual fidelity, structural preservation, target localization, and motion consistency. Experiments demonstrate that CrossScope consistently outperforms strong surgical video generation baselines, validating the importance of role-aware evidence routing for multi-observer visual world modeling.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.