2602.11024v1 Feb 11, 2026 cs.CV

밀집된 수술 기구 계수를 위한 시선 연쇄 기반 공간 추론

Chain-of-Look Spatial Reasoning for Dense Surgical Instrument Counting

Rishikesh Bhyri
Rishikesh Bhyri
Citations: 1
h-index: 1
Brian R Quaranto
Brian R Quaranto
Citations: 379
h-index: 11
Philip J. Seger
Philip J. Seger
Citations: 82
h-index: 2
Kaity Tung
Kaity Tung
Citations: 8
h-index: 1
B. Fox
B. Fox
Citations: 3
h-index: 1
Gene Yang
Gene Yang
Citations: 36
h-index: 1
Steven D. Schwaitzberg
Steven D. Schwaitzberg
Citations: 136
h-index: 5
Junsong Yuan
Junsong Yuan
Citations: 15
h-index: 2
Nan Xi
Nan Xi
Citations: 17
h-index: 2
Peter C W Kim
Peter C W Kim
Citations: 45
h-index: 3

수술실에서 수술 중 환자 안전을 보장하기 위한 중요한 전제 조건은 정확한 수술 기구 계수입니다. 최근의 거대 시각-언어 모델과 에이전트 AI의 발전에도 불구하고, 특히 기구가 밀집되어 있는 경우, 이러한 기구를 정확하게 계수하는 것은 여전히 매우 어려운 과제입니다. 이러한 문제를 해결하기 위해, 우리는 기존의 객체 탐지 방식이 가지는 순서 제약 없이, 구조화된 시각적 연쇄를 통해 인간의 계수 과정을 모방하는 새로운 시각 추론 프레임워크인 Chain-of-Look을 제안합니다. 이 시각적 연쇄는 모델이 일관된 공간 경로를 따라 계수를 수행하도록 안내하여 복잡한 장면에서의 정확도를 향상시킵니다. 또한, 밀집된 수술 기구의 물리적 제약을 명시적으로 모델링하는 이웃 손실 함수를 도입하여 시각적 연쇄의 물리적 타당성을 더욱 강화했습니다. 또한, 1,464개의 고밀도 수술 기구 이미지를 포함하는 새로운 데이터셋인 SurgCount-HD를 공개합니다. 광범위한 실험 결과, 제안하는 방법이 기존의 최첨단 계수 방법(예: CountGD, REC) 및 다중 모드 대규모 언어 모델(예: Qwen, ChatGPT)보다 밀집된 수술 기구 계수라는 어려운 과제에서 우수한 성능을 보이는 것을 확인했습니다.

Original Abstract

Accurate counting of surgical instruments in Operating Rooms (OR) is a critical prerequisite for ensuring patient safety during surgery. Despite recent progress of large visual-language models and agentic AI, accurately counting such instruments remains highly challenging, particularly in dense scenarios where instruments are tightly clustered. To address this problem, we introduce Chain-of-Look, a novel visual reasoning framework that mimics the sequential human counting process by enforcing a structured visual chain, rather than relying on classic object detection which is unordered. This visual chain guides the model to count along a coherent spatial trajectory, improving accuracy in complex scenes. To further enforce the physical plausibility of the visual chain, we introduce the neighboring loss function, which explicitly models the spatial constraints inherent to densely packed surgical instruments. We also present SurgCount-HD, a new dataset comprising 1,464 high-density surgical instrument images. Extensive experiments demonstrate that our method outperforms state-of-the-art approaches for counting (e.g., CountGD, REC) as well as Multimodality Large Language Models (e.g., Qwen, ChatGPT) in the challenging task of dense surgical instrument counting.

1 Citations
0 Influential
5.5 Altmetric
28.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!