2606.26734v1 Jun 25, 2026 cs.CV

강건한 양파: 노이즈 환경에서 개방형 어휘 객체 검출기의 성능 분석

Robust Onion: Peeling Open Vocab Object Detectors Under Noise

Y. Rawat
Y. Rawat
Citations: 2,474
h-index: 24
P. Pathak
P. Pathak
Citations: 52
h-index: 4
Mukilan Karuppasamy
Mukilan Karuppasamy
Citations: 8
h-index: 2
Aaditya Baranwal
Aaditya Baranwal
Citations: 3
h-index: 1
Shruti Vyas
Shruti Vyas
Citations: 565
h-index: 13

개방형 어휘 객체 검출기(OV-OD)의 복잡한 구조로 인해 실제 환경에서의 노이즈가 OV-OD에 미치는 영향은 제대로 이해되지 못하고 있습니다. 본 연구에서는 '강건한 양파(Robust Onion)'라는 이름으로, 통제된 인공적인 시각적 왜곡을 사용하여 OV-OD를 층별로 분석하는 종합적인 연구를 수행했습니다. 이를 통해 모델의 강건성이 저하되는 방식, 이유 및 위치를 파악하고, 특징 지도(feature map)의 붕괴 현상을 체계적으로 분석했습니다. 연구 결과, 유사한 시각적 기반 구조를 가진 모델들은 유사한 층에서 발생하는 유사한 특징 지도 붕괴 현상으로 인해 비슷한 수준의 강건성을 보이는 것으로 나타났습니다. 반면, 사전 학습 전략, 아키텍처상의 미묘한 차이, 그리고 캡션 기반의 지도 학습은 강건성에 큰 영향을 미치지 않는 것으로 확인되었습니다. 강건성은 주로 이미지 도메인에 의해 결정되며, 어노테이션과는 직접적인 관련이 적습니다. 이러한 이유로 인해 COCO 및 LVIS와 같은 데이터셋에서 유사한 수준의 강건성 변화가 나타나며, ODinW-13과 같이 큰 크기의 개별 객체가 많은 데이터셋은 인위적으로 과장된 강건성을 보이는 것처럼 보일 수 있습니다. 마지막으로, 저희는 개발한 경량화된 플러그 앤 플레이 NN & TK0 방법을 사용하여 실제 환경인 BDD100K, WiderFace 및 VisDRONE 데이터셋에서 모델의 강건성을 향상시켰습니다. 이 방법은 전체 학습 방식에 비해 96배 적은 학습 가능한 파라미터를 사용합니다. 또한, 기존 연구들의 강건성 관련 관찰 결과에 대한 설명도 제공합니다.

Original Abstract

The impact of real-world noise on Open Vocabulary Object Detectors (OV-ODs) remains poorly understood due to their architectural complexity. We present our comprehensive analysis Robust Onion, an empirical study that uses controlled synthetic visual degradations to peel OV-ODs layer-by-layer, revealing how, why, and where robustness degrades, systematically analyzing feature collapse. Our findings reveal that models with similar vision backbones exhibit comparable robustness, driven by similar feature collapse at similar layers, while factors such as pretraining strategy, architectural nuances, and caption supervision contribute little. Robustness is primarily governed by the image domain rather than annotations, explaining the similar robustness impact on COCO and LVIS, and why datasets like ODinW-13 can give an impression of inflated robustness due to large, isolated objects. Finally, we validate our insights by improving robustness on real-world BDD100K, WiderFace, and VisDRONE via our lightweight plug-and-play NN & TK0 approach, using 96x fewer trainable parameters than end-to-end training. We also explain the prior works' robustness observations.

1 Citations
0 Influential
12 Altmetric
61.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!