2607.25626v1 Jul 28, 2026 cs.AI

중국어 음성 생성 및 인지 과정에서 뇌파(EEG)를 기반으로 한 문장 변환을 위한 텍스트-오디오 통합 정렬 연구

Joint Text-Audio Alignment for EEG-to-Text Decoding in Chinese Speech Production and Perception

Xinxin Zhu
Xinxin Zhu
Citations: 1,357
h-index: 15
Tian Zheng
Tian Zheng
Citations: 6
h-index: 2
Feng Tian
Feng Tian
Citations: 23
h-index: 3
Xurong Xie
Xurong Xie
Citations: 11
h-index: 2
Xiaolan Peng
Xiaolan Peng
Citations: 309
h-index: 11

뇌파(EEG)로부터 직접적으로 음성 정보를 텍스트로 변환하는 것은 심각한 언어 및 운동 장애가 있는 사람들을 위한 잠재적인 비침습적 신경 통신 경로를 제공할 수 있습니다. 침습적인 방법인 전기피질 기록법에 비해 EEG는 더 안전하고 광범위하게 적용 가능하지만, 변환하기 훨씬 어렵습니다. 특히 중국어 문장 변환은 수천 개의 문자를 포함하는 고차원 출력 공간, 개인 간의 심각한 변동성, 그리고 텍스트 정렬을 위한 낮은 신호 대 잡음비로 인해 어려움이 더욱 가중됩니다. 기존 방법들은 단일 감독 신호 축(즉, 텍스트 의미 또는 오디오 음향 특징)에 의존하지만, 문장 수준의 구별력과 큰 어휘 크기의 중국어 변환에 필요한 세밀한 시간 분해능을 동시에 만족시킬 수 없습니다. 본 연구에서는 뇌파와 두 가지 축을 함께 정렬하는 새로운 효율적인 프레임워크인 EEGAlign을 제안합니다. 이는 BGE-M3 텍스트 임베딩을 사용한 텍스트 정렬과 wav2vec~2.0 음성 특징을 이용한 오디오 정렬을 대조 학습을 통해 수행하고, 이어 CTC(Connectionist Temporal Classification)를 사용하여 문자 시퀀스를 변환합니다. 중국EEG-2 데이터셋에서 EEGAlign은 최고 수준의 문장 분류 성능을 달성했으며, 낭독 뇌파 데이터에서 최대 82.37%의 Top-1 정확도, 수동 청취 뇌파 데이터에서 41.43%의 정확도를 보였습니다 (101개의 후보 문장 중). 제거 실험 결과, 두 가지 정렬 축은 서로 보완적이며, 함께 사용할 때 개별적으로 사용하는 것보다 일관되게 더 나은 성능을 나타냅니다. 본 연구는 공개 음성 생성 과정에서 비침습적인 뇌파를 사용하여 큰 어휘 크기의 중국어 문장을 변환하는 첫 번째 연구이며, 상대적으로 큰 후보 문장 집합에서도 강력한 분류 성능을 달성했습니다.

Original Abstract

Decoding speech information directly from scalp electroencephalography (EEG) into text provides a potential non-invasive neural communication pathway for individuals with severe speech and motor impairments. Compared with invasive approaches such as electrocorticography, EEG is safer and more widely deployable, yet substantially more challenging to decode.This challenge is exacerbated for Chinese sentence decoding, which must handle a high-dimensional output space with thousands of characters, severe inter-subject variability, and low signal-to-noise ratios for text alignment.Existing methods commit to a single supervisory axis---either text semantics or audio acoustic features---yet neither can simultaneously satisfy the demands of sentence-level discriminability and fine-grained temporal resolution required for large-vocabulary Chinese decoding. We introduce EEGAlign, a novel parameter-efficient framework that jointly aligns EEG with two axes---text alignment with BGE-M3 text embeddings and audio alignment with wav2vec~2.0 speech features via contrastive learning followed by CTC character-sequence decoding. On ChineseEEG-2 data, EEGAlign yields state-of-the-art closed-set sentence classification performance, reaching up to 82.37% Top-1 accuracy on Reading Aloud EEG and 41.43% on Passive Listening EEG out of 101 candidates. Ablation studies show that the two alignment axes are highly complementary: combining them yields consistently better performance than either alone. To the best of our knowledge, this is the first study on decoding large-vocabulary Chinese sentences from non-invasive EEG during overt speech production, and achieving strong classification performance with relatively large closed-set candidate-sentence setting.

0 Citations
0 Influential
7.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!