그래프 기반 다중 인스턴스 학습을 통한 암묵적 및 명시적 관계 편향 통합: 피부 병변 진단 사례 연구
Integrating Implicit and Explicit Relational Biases through Graph-Based Multiple Instance Learning: A Case Study in Skin Lesion Diagnosis
관계적 유도 편향은 데이터 간의 구조적 의존성을 파악하는 데 필수적입니다. 본 연구는 이미지 분류를 위한 이중 수준의 관계 프레임워크를 조사하며, 암묵적인 표현 학습과 명시적인 구조 모델링 간의 격차를 해소합니다. 먼저 EfficientNetB3 아키텍처를 사용하여 기준 성능을 확립했습니다. 표준 컨볼루션 편향에서 벗어나 패치 기반 전략을 채택하여, 자기 지도 방식으로 재구성을 통해 패치 간의 암묵적인 관계를 학습하기 위해 컨볼루션 마스크 오토인코더를 사용합니다. 그런 다음 이 접근 방식을 확장하여 명시적인 관계 모델링을 통합하고, 학습된 임베딩을 그리드 기반, 랜덤 및 k-최근접 이웃 구조를 포함한 다양한 그래프 토폴로지로 구성합니다. ISIC-2018 및 ISIC-2019 피부 병변 진단 벤치마크에 대한 실험 결과는 암묵적인 패치 기반 모델링과 명시적인 그래프 기반 메시지 전달을 결합하는 것이 최상의 성능을 제공한다는 것을 보여줍니다. ISIC-2018 테스트 세트에서 기준 모델은 76.17%의 균형 정확도를 달성했으며, 암묵적인 패치 기반 관계 모델링을 통해 이 수치가 77.12%로 향상되었습니다. 완전히 통합된 그리드 구조를 가진 그래프 어텐션 네트워크는 성능을 더욱 향상시켜 79.27%에 도달했습니다. 마찬가지로 ISIC-2019에서는 암묵적인 접근 방식이 59.84%의 균형 정확도를 달성하는 반면, 암묵적 및 명시적 모델링의 조합은 60.67%를 달성합니다.
Relational inductive biases are essential for capturing structural dependencies among data. This study investigates a dual-level relational framework for image classification, bridging the gap between implicit representation learning and explicit structural modelling. We begin by establishing a baseline using an EfficientNetB3 architecture. To move beyond standard convolutional biases, we adopt a patch-based strategy, employing a convolutional masked autoencoder to learn implicit inter-patch relationships through self-supervised reconstruction. We then extend this approach by incorporating explicit relational modelling, organizing the learned embeddings into various graph topologies, including grid-based, random, and k-nearest neighbour structures. Experimental results on the ISIC-2018 and ISIC-2019 skin lesion diagnosis benchmarks show that combining implicit inter-patch modelling with explicit graph-based message passing yields the best performance. On the ISIC-2018 test set, the baseline model achieves a balanced accuracy of 76.17%, which improves to 77.12% with implicit patch-based relational modelling. The fully integrated grid-structured Graph Attention Network further increases performance to 79.27%. Similarly, on ISIC-2019, the implicit approach reaches 59.84% balanced accuracy, while the combination of implicit and explicit modelling yields 60.67%.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.