SIRR-LMM: 대규모 다중 모드 모델을 이용한 단일 이미지 반사 제거
SIRR-LMM: Single-image Reflection Removal via Large Multimodal Model
유리 표면은 반사된 빛과 투과된 빛의 복잡한 상호작용을 만들어내기 때문에, 단일 이미지 반사 제거(SIRR)는 어려운 문제입니다. 기존 데이터셋은 합성 데이터의 경우 물리적 현실성이 부족하거나, 실제 데이터의 경우 데이터 규모가 충분하지 않은 문제가 있습니다. 우리는 3D 유리 모델을 실제 배경 이미지에 적용하여 물리적으로 정확한 반사 시나리오를 생성하는 합성 데이터셋 생성 프레임워크를 소개합니다. 이 프레임워크는 다양한 유리 특성, 카메라 설정 및 후처리 효과를 포함합니다. 대규모 다중 모드 모델(LMM)의 기능을 활용하기 위해, 이미지 레이어를 하나의 복합 입력으로 연결하고, 공동 캡셔닝을 적용하며, 전체 파라미터를 학습하는 대신 작업에 특화된 LoRA를 사용하여 모델을 미세 조정합니다. 이를 통해 저희의 접근 방식은 최첨단 방법과 비교하여 향상된 반사 제거 및 분리 성능을 달성할 수 있습니다.
Glass surfaces create complex interactions of reflected and transmitted light, making single-image reflection removal (SIRR) challenging. Existing datasets suffer from limited physical realism in synthetic data or insufficient scale in real captures. We introduce a synthetic dataset generation framework that path-traces 3D glass models over real background imagery to create physically accurate reflection scenarios with varied glass properties, camera settings, and post-processing effects. To leverage the capabilities of Large Multimodal Model (LMM), we concatenate the image layers into a single composite input, apply joint captioning, and fine-tune the model using task-specific LoRA rather than full-parameter training. This enables our approach to achieve improved reflection removal and separation performance compared to state-of-the-art methods.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.