2607.11801v1 Jul 13, 2026 cs.SD

대규모 오디오-언어 모델에서 음향 인식을 위한 인코더 측 신경 세포 식별 및 증폭

Encoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language Models

Anqi Cheng
Anqi Cheng
Citations: 1
h-index: 1
Chih-Kai Yang
Chih-Kai Yang
Citations: 389
h-index: 12
Ke-Han Lu
Ke-Han Lu
Citations: 518
h-index: 13
Hung-yi Lee
Hung-yi Lee
Citations: 124
h-index: 4
Yu-Han Huang
Yu-Han Huang
Citations: 45
h-index: 3

대규모 오디오-언어 모델(LALM)은 발화 내용에 대한 강력한 성능을 보이는 반면, 화자의 감정과 같은 미세하고 비의미적인 음성 속성에 대해서는 종종 낮은 성능을 보입니다. 이러한 문제를 재학습 없이 개선하기 위해서는 효과적인 추론 시간 개입이 필요하며, 기존 방법 대부분은 오디오 인코더 이후에만 작동하고 비교적 거친 수준에서 작동합니다. 본 논문에서는 음향 정보가 먼저 파형으로부터 추출되는 인코더 자체, 특히 개별 신경 세포 수준에서 연구가 미흡하다는 점에 주목했습니다. 우리는 IAAN(Identifying and Amplifying Acoustic Neurons)이라는 새로운 방법을 제안하며, 이는 학습 및 레이블 없이 오디오 인코더의 각 순방향 신경 세포를 평가합니다. IAAN은 실제 파형에서의 활성화와 실제 오디오의 음향 정보를 포함하지 않는 노이즈 참조에 대한 활성화를 비교하여 각 신경 세포에 점수를 부여합니다. 추론 시, IAAN은 가장 높은 점수를 받은 소수의 신경 세포를 증폭시킵니다. 실험 결과, IAAN은 Audio-Flamingo-3에서 평균 정확도를 25.7점, Qwen2.5-Omni에서 21.4점, Kimi-Audio에서 9.7점 향상시켰습니다. 또한 음향 정보를 우선적으로 고려하도록 명시적으로 미세 조정된 모델의 성능도 개선했습니다. 통제 비교 실험 결과, 인코더 위치와 신경 세포 수준의 선택성이 모두 이러한 성능 향상에 필수적임을 확인했습니다. 디코딩 단계 이후 또는 언어 모델 내에서의 개입은 거의 또는 전혀 성능 향상을 가져오지 않거나 오히려 정확도를 저하시켰습니다. 또한 개선 효과는 증폭되는 특정 신경 세포에 따라 달라지며, 단순히 신경 세포의 수에만 의존하지 않는다는 점을 확인하여 IAAN의 음향 점수가 중요한 신경 세포를 식별하는 데 성공했음을 입증했습니다. 이러한 결과는 오디오 인코더 내에서 작고 정밀하게 설계된 개입이 LALM의 음향 이해력을 강화하는 효과적이고 활용도가 낮은 방법임을 보여주며, 신경 세포 수준에서 인코더에 접근하여 음향 인식을 개선하는 추론 시간 방법을 위한 새로운 방향을 제시합니다.

Original Abstract

Large audio-language models (LALMs) often underperform on fine-grained, non-semantic attributes of speech, such as a speaker's emotion, despite strong performance on speech content. Improving this without the cost of retraining calls for an effective inference-time intervention, yet most existing methods intervene only after the audio encoder and operate at a relatively coarse granularity. The encoder itself, where acoustic information is first extracted from the waveform, remains largely unexplored, especially at the level of individual neurons. We introduce IAAN, Identifying and Amplifying Acoustic Neurons, a training-free and label-free method that scores each feed-forward neuron in the audio encoder by contrasting its activation on the real waveform with that on a noise reference lacking the real audio's acoustic information. IAAN then amplifies a small set of the highest-scoring neurons at inference. Across ten non-semantic speech attributes, IAAN improves average accuracy by 25.7 points on Audio-Flamingo-3, 21.4 on Qwen2.5-Omni, and 9.7 on Kimi-Audio. It also improves a model already explicitly fine-tuned to prioritize acoustic evidence. In controlled comparisons, both the encoder locus and neuron-level selectivity prove necessary for this gain. Intervening after the encoder, at the decoding side or inside the language model, yields little to no improvement, or even deteriorates accuracy. The improvement also depends on which specific neurons are amplified, not merely on their number, confirming that IAAN's acoustic score succeeds in identifying the neurons that matter. These results show that a small, precisely targeted intervention inside the audio encoder is an effective and largely untapped way to strengthen the acoustic understanding of LALMs, opening a new direction for inference-time methods that improve acoustic perception through neuron-level access to the encoder.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!