2606.31834v1 Jun 30, 2026 cs.CV

실시간 소스 없는 객체 탐지

Real-Time Source-Free Object Detection

Vineeth N. Balasubramanian
Vineeth N. Balasubramanian
Citations: 390
h-index: 8
Varun Gopal
Varun Gopal
Citations: 0
h-index: 0
Poornima Jain
Poornima Jain
Citations: 27
h-index: 4
Muhammad Haris Khan
Muhammad Haris Khan
Citations: 37
h-index: 4
S. V. Rebbapragada
S. V. Rebbapragada
Citations: 19
h-index: 2

자율 주행, 감시 및 로봇 공학을 위한 실세계 검출기는 엄격한 지연 시간 및 메모리 제약 조건 하에서 도메인 변화를 처리해야 하지만, 기존의 소스 없는 객체 탐지(SFOD) 방법은 정확도를 우선시하는 무거운 아키텍처에 의존합니다. 우리는 이러한 절충이 불필요하다는 것을 보여줍니다. YOLOv10을 기반으로 구축된 NMS-free 이중 헤드 검출기를 사용하여, 이전 최고 수준의 적응 정확도를 달성하면서 더 빠르고 가벼운 모델을 만들었습니다. 이중 헤드 검출기에 일반적인 평균-교사(mean-teacher) 자기 학습을 직접 적용하면 두 가지 주요 요인으로 인해 최적 이하의 적응 성능이 나타납니다. 첫째, 단일 헤드를 사용하거나 양쪽 헤드의 높은 신뢰도 예측을 직접 결합하는 것과 같은 간단한 유사 레이블 생성 전략은 도메인 변화 시에 최적이 아닌 감독을 제공합니다. 우리는 정밀도를 유지하고 누락된 객체를 복구하기 위해 하나 대 하나(O2O) 및 하나 대 다수(O2M) 헤드 예측을 선택적으로 허용하는 DHF(이중 헤드 유사 레이블 융합)를 제안합니다. 둘째, 도메인 변화는 멀티 스케일 특징의 구별력을 저하시킨다는 것을 관찰했습니다. 우리는 객체 탐지에 대한 변동 및 공분산 제약을 멀티 스케일 특징 맵에 적용하여 이를 완화하는 MARD(멀티 스케일 적응 표현 다양화) 손실을 사용합니다. 두 모듈 모두 학습 시간에만 적용되며, 추론에는 영향을 미치지 않습니다. 다양한 도메인 변화 벤치마크에서, 저희 방법인 RT-SFOD는 기존 최고 수준의 SFOD 방법에 비해 1.4~3.5%의 mAP 향상, 1.3배 더 빠른 처리 속도, 그리고 약 2배 적은 파라미터를 제공하며, 속도-정확도-모델 크기 간의 균형을 더욱 발전시켰습니다. 저희는 YOLOv10을 사용하여 주요 결과를 보고했으며, 추가적인 YOLO 및 DETR 기반 이중 헤드 검출기를 통해 일반화 성능을 입증했습니다. 코드는 다음 주소에서 확인할 수 있습니다: https://github.com/Sairam13001/RT-SFOD/

Original Abstract

Real-world detectors for autonomous driving, surveillance, and robotics must handle domain-shifts under strict latency and memory constraints, yet existing source-free object detection (SFOD) methods rely on heavyweight architectures that prioritize accuracy alone. We show this trade-off is unnecessary: building on YOLOv10, an NMS-free dual-head detector, we achieve state-of-the-art adaptation accuracy while being faster and more compact. We observe that directly applying vanilla mean-teacher self-training to dual-head detectors leads to suboptimal adaptation performance due to two key factors. First, simple pseudo-label generation strategies, such as using a single head or directly combining high-confidence predictions from both heads, yield suboptimal supervision under domain-shift. We propose DHF (Dual-Head Pseudo-Label Fusion) which selectively admits one-to-one (O2O) and one-to-many (O2M) head predictions, preserving precision and recovering missed objects. Second, we observe domain-shift collapses multi-scale feature discriminability. We propose the use of our MARD (Multi-scale Adaptive Representation Diversification) loss which mitigates this by enforcing detection-aware variance and covariance constraints on multi-scale feature maps. Both modules are training-time only, leaving inference unchanged. Across domain-shift benchmarks, our method, RT-SFOD yields 1.4 to 3.5\% mAP gains, 1.3$\times$ higher throughput, with $\sim$2$\times$ fewer parameters than prior state-of-the-art SFOD methods, thus advancing the Pareto frontier of the speed-accuracy-model size trade-off. We report main results with YOLOv10, and demonstrate generalizability with additional YOLO- and DETR-based dual-head detectors. Code is available here: https://github.com/Sairam13001/RT-SFOD/

0 Citations
0 Influential
24 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!