심층 신경망 분류기를 위한 정규화된 조건부 상호 정보 기반 대체 손실
Normalized Conditional Mutual Information Surrogate Loss for Deep Neural Classifiers
본 논문에서는 심층 신경망(DNN) 기반 분류기 학습을 위한 기존의 교차 엔트로피(CE) 손실 함수를 대체할 수 있는 새로운 정보 이론 기반의 대체 손실 함수인 정규화된 조건부 상호 정보(NCMI)를 제안합니다. 먼저, 모델의 NCMI 값과 정확도의 역 상관 관계를 확인합니다. 이러한 통찰력을 바탕으로, NCMI를 효율적으로 최소화하는 교대 알고리즘을 도입합니다. 이미지 인식 및 전체 슬라이드 이미징(WSI) 분류 벤치마크에서 NCMI를 사용하여 학습된 모델은 상당한 수준으로 기존 최고 성능의 손실 함수를 능가하며, CE 손실 함수와 비교했을 때 계산 비용은 거의 동일합니다. 특히, ImageNet 데이터셋에서 ResNet-50 모델을 사용할 때 NCMI는 CE 손실 함수에 비해 2.77%의 top-1 정확도 향상을 보입니다. CAMELYON-17 데이터셋에서는 NCMI를 CE 손실 함수로 대체함으로써, 최고 성능의 기준 모델보다 macro-F1 점수가 8.6% 향상되었습니다. 이러한 성능 향상은 다양한 아키텍처와 배치 크기에서 일관적으로 나타나며, 이는 NCMI가 CE 손실 함수에 대한 실용적이고 경쟁력 있는 대안임을 시사합니다.
In this paper, we propose a novel information theoretic surrogate loss; normalized conditional mutual information (NCMI); as a drop in alternative to the de facto cross-entropy (CE) for training deep neural network (DNN) based classifiers. We first observe that the model's NCMI is inversely proportional to its accuracy. Building on this insight, we introduce an alternating algorithm to efficiently minimize the NCMI. Across image recognition and whole-slide imaging (WSI) subtyping benchmarks, NCMI-trained models surpass state of the art losses by substantial margins at a computational cost comparable to that of CE. Notably, on ImageNet, NCMI yields a 2.77% top-1 accuracy improvement with ResNet-50 comparing to the CE; on CAMELYON-17, replacing CE with NCMI improves the macro-F1 by 8.6% over the strongest baseline. Gains are consistent across various architectures and batch sizes, suggesting that NCMI is a practical and competitive alternative to CE.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.