2607.27274v1 Jul 29, 2026 cs.LG

뇌파 기반 질병 진단의 재고찰: 개체 표현 학습과 피험자 수준의 지도 학습 분리

Rethinking EEG-Based Disease Diagnosis: Decoupling Instance Representation Learning from Subject-Level Supervision

Yuhao Sun
Yuhao Sun
Citations: 823
h-index: 3
Zhiyuan Ma
Zhiyuan Ma
Citations: 0
h-index: 0
Xinche Zhang
Xinche Zhang
Citations: 11
h-index: 2
Sen Song
Sen Song
Citations: 439
h-index: 2
Zeyuan Li
Zeyuan Li
Citations: 56
h-index: 4
Xinke Shen
Xinke Shen
Citations: 11
h-index: 3
Zhiyi Lu
Zhiyi Lu
Citations: 0
h-index: 0
Jiacheng Hao
Jiacheng Hao
Citations: 9
h-index: 2
Youlang Du
Youlang Du
Citations: 0
h-index: 0
Zhengxu Jiang
Zhengxu Jiang
Citations: 0
h-index: 0

뇌파 기반 질병 진단은 각 피험자에 대해 하나의 예측을 필요로 하지만, 일반적인 방법에서는 기록을 짧은 단위로 나누고, 모든 단위에 해당 피험자의 레이블을 적용하여 개체 수준 분류기를 훈련합니다. 이는 모든 개체가 동일하게 신뢰할 수 있는 진단 증거를 제공한다고 가정합니다. 다중 개체 학습(MIL)은 상속된 레이블 문제를 해결하기 위해 각 피험자를 '가방'으로 취급하지만, 뇌파 데이터셋은 일반적으로 개체보다 피험자의 수가 훨씬 적어, end-to-end MIL을 통해 학습되는 표현의 품질을 제한할 수 있습니다. 본 연구에서는 개체 표현 학습과 피험자 수준 지도 학습을 분리하는 두 단계 프레임워크인 BridgeMIL을 제안합니다. 1단계에서는 시간적으로 인접한 창을 정렬하고, 각 피험자 내에서 독립적으로 샘플링된 부분 가방을 사용하여 개체 레이블 없이 인코더를 사전 훈련합니다. 분산 및 공분산 정규화는 붕괴를 방지하고 중복성을 줄이며, 부정적인 쌍을 사용하지 않습니다. 2단계에서는 학습된 인코더를 어텐션 기반 MIL 집계기로 이전하고, 피험자 예측에만 지도 학습을 적용하며, 특징 보존을 통해 표현의 변화를 제한합니다. 세 개의 뇌파 질병 데이터셋과 다섯 가지 대표적인 모델 구조에서 BridgeMIL은 15개의 데이터셋-모델 조합 중 14개에서 가장 높은 평균 정확도를 달성했으며, 전반적인 평균 정확도는 76.57%로, 가장 강력한 기준 모델보다 4.28% 포인트 더 높습니다. 추가 분석 결과, 개체별 레이블의 신뢰성이 개체마다 크게 다르며, 피험자 부족에 대한 성능 민감도가 개체 부족보다 더 크다는 것을 확인했습니다. 또한, 진단 클래스 간의 분리가 개선된, 뚜렷한 피험자 기반 클러스터를 갖는 보다 구조화된 표현 공간이 형성되었습니다. 이러한 결과들은 뇌파 데이터를 활용하여 풍부한 정보를 얻으면서도, 개체 수준 예측 목표에 맞게 지도 학습을 수행하는 것의 중요성을 강조합니다.

Original Abstract

EEG-based disease diagnosis requires one prediction per subject, yet common pipelines segment recordings into short instances, inherit the subject label for every instance, and train instance-level classifiers. This assumes that all instances provide equally reliable diagnostic evidence. Multiple instance learning (MIL) avoids inherited labels by treating each subject as a bag. However, EEG datasets contain far fewer subjects than instances, which can limit the quality of the representations learned by end-to-end MIL. We propose BridgeMIL, a two-stage framework that decouples instance representation learning from subject-level supervision. Stage 1 pretrains the encoder without inherited instance labels by aligning temporally nearby windows and independently sampled within-subject sub-bags. Variance and covariance regularization prevent collapse and reduce redundancy without negative pairs. Stage 2 transfers the encoder to an attention-based MIL aggregator, applies supervision only to subject predictions, and limits representation drift through feature retention. Across three EEG disease datasets and five representative backbones, BridgeMIL attains the highest mean accuracy in 14 of 15 dataset-backbone settings and an overall mean accuracy of 76.57%, 4.28 percentage points higher than the strongest baseline. Further analyses reveal substantial variation in inherited-label reliability across instances, greater performance sensitivity to subject scarcity than to instance scarcity, and a more structured representation space with distinct subject-wise clusters and improved separation between diagnostic classes. Together, these findings underscore the importance of aligning supervision with the subject-level prediction objective while learning from abundant EEG instances without assigning disease labels to individual instances.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!