2608.05109v1 Aug 05, 2026 eess.IV

AI 기반 단일 촬영 구조광 깊이 복원을 통한 실시간 복강경 수술 안내

AI-based single-shot structured-light depth reconstruction for real-time laparoscopic surgical guidance

M. Kam
M. Kam
Citations: 547
h-index: 9
J. Opfermann
J. Opfermann
Citations: 1,928
h-index: 19
Wayne Wonseok Rodgers
Wayne Wonseok Rodgers
Citations: 1
h-index: 1
Xiangyi Le
Xiangyi Le
Citations: 0
h-index: 0
Seonghoon Jang
Seonghoon Jang
Citations: 1,033
h-index: 14
Axel Krieger
Axel Krieger
Citations: 393
h-index: 5
Shuwen Wei
Shuwen Wei
Citations: 706
h-index: 9
Jin U. Kang
Jin U. Kang
Citations: 517
h-index: 8

배경 및 중요성: 정밀한 수술 중 깊이 인식은 자율 또는 반자율 복강 로봇 수술에 매우 중요하다. 기존의 간섭 패턴 투영 프로파일로메트리는 밀리미터 수준의 정확도를 달성할 수 있지만, 종종 다중 촬영, 디지털 마이크로미러 장치 투영 및 투사기-카메라 동기화가 필요하여 소형 복강경 시스템에 통합하기 어렵다. 목표: 수동 LED 조명을 사용하는 이진 마스크와 맞춤형 U-Net 깊이 모듈을 갖춘 VQ-VAE 사전 지식을 활용하여 동기화가 불필요한 단일 촬영 심도 감지 플랫폼을 개발하는 것을 목표로 한다. 방법: 소형 투사 모듈을 이중 채널 복강경의 한 채널에 결합하고, 다른 채널은 간섭 패턴으로 조명된 표적을 이미징한다. Zivid 3D 카메라는 722개의 쌍을 이루는 팬텀 이미지에 대한 참조 깊이 데이터를 획득했으며, 이 데이터는 지도 학습 및 평가를 위해 SSLE (Structured-light Single-Exposure) 이미지 프레임으로 재투영되었다. VQ-VAE는 각 입력 데이터를 이산적인 잠재 표현으로 인코딩하고, 잠재 공간 U-Net은 별도의 마스크 예측 브랜치 없이 깊이를 예측한다. 결과: 고정된 학습/검증/테스트 세트를 사용하여 제안된 모델은 MAE 3.70 mm, AbsRel 0.0326, delta=1.1 정확도 0.962, delta=1.1^2 정확도 0.970을 달성했다. 이 모델은 이중 U-Net MaskNet + DepthNet 기준 성능보다 낮은 MAE를 보였으며, 상용 단안 깊이 모델보다 MAE, AbsRel 및 임계값 정확도 측면에서 더 우수한 성능을 나타냈다. 파이프라인은 NVIDIA A100 GPU에서 301개의 연속 프레임에 대해 26.0 Hz로 작동했다. 결론: 수동 LED 조명을 사용하는 이진 패턴 플랫폼과 잠재 공간 깊이 복원은 동기화가 불필요하고 비디오 속도의 내시경 심도 추정을 가능하게 한다. 결과는 명시적인 분할 단계 없이 Zivid 참조 팬텀 재구성을 보여주며, 데이터 세트 크기와 SSLE-Zivid 교정 정확도가 중요하다는 점을 강조한다.

Original Abstract

Significance. Accurate intraoperative depth perception is important for autonomous and semi-autonomous robotic laparoscopic surgery. Conventional fringe projection profilometry can achieve millimeter-scale accuracy but often requires multi-shot acquisition, digital-micromirror-device projection, and projector-camera synchronization, complicating integration into compact laparoscopic systems. Aim. To develop a synchronization-free, single-shot depth-sensing platform using a passive LED-illuminated binary mask and a VQ-VAE prior with a custom U-Net depth head. Approach. A compact projection module was coupled to one channel of a dual-channel laparoscope, while the second channel imaged the fringe-illuminated target. A Zivid 3D camera acquired reference depth for 722 paired phantom images. Zivid depth maps were reprojected into the SSLE image frame for supervised training and evaluation. The VQ-VAE encoded each input into a discrete latent representation, and a latent-space U-Net predicted depth without a separate mask-prediction branch. Results. Using a fixed train/validation/test split, the proposed model achieved an MAE of 3.70 mm, AbsRel of 0.0326, delta=1.1 accuracy of 0.962, and delta=1.1^2 accuracy of 0.970. It achieved lower MAE than the dual U-Net MaskNet + DepthNet baseline and outperformed off-the-shelf monocular depth models in MAE, AbsRel, and threshold accuracy. The pipeline operated at 26.0 Hz over 301 consecutive frames on an NVIDIA A100 GPU. Conclusions. The LED-illuminated binary-pattern platform with latent-space depth reconstruction enables synchronization-free, video-rate endoscopic depth estimation. Results demonstrate Zivid-referenced phantom reconstruction without an explicit segmentation stage, while emphasizing the importance of dataset size and SSLE-Zivid calibration accuracy.

0 Citations
0 Influential
9.5 Altmetric
47.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!