타겟 언어 특성 강화: SAE 기반 다국어 추론 제어
Strengthening Target-Language Features: SAE-Based Steering for Multilingual Inference
다국어 대규모 언어 모델은 언어별 성능 차이가 크며, 기존 적응 방법들은 종종 파라미터 업데이트와 상당한 양의 다국어 학습 데이터를 필요로 합니다. 본 논문에서는 사전 훈련된 희소 자동 인코더(sparse autoencoders, SAE)를 사용하여 타겟 언어 관련 특징을 식별하고 강화하는 추론 시간 기반 다국어 제어 방법을 제안합니다. 다국어 병렬 문장을 활용하여, 각 타겟 언어와 관련된 특정 레이어의 특징을 선택하기 위해 다양한 언어에서의 SAE 활성화를 비교 분석합니다. 이러한 특징들은 제어 신호로 디코딩되어 추가적인 훈련 없이 모델의 숨겨진 상태에 주입됩니다. Gemma-3-12B-it 모델을 사용한 실험 결과, XCOPA에서 평균 정확도가 10.9%p, XNLI에서 5.3%p, MGSM에서 1.9%p 향상되었습니다.
Multilingual large language models exhibit substantial performance differences across languages, while existing adaptation methods often require parameter updates and considerable multilingual training data. We propose an inference-time multilingual steering method that uses pretrained sparse autoencoders to identify and strengthen target-language-related features. Using multilingual parallel sentences, we compare SAE activations across languages and select a small number of layer-specific features associated with each target language. These features are decoded into steering signals and injected into the model's hidden states without additional training. Experiments with Gemma-3-12B-it show average accuracy improvements of 10.9 percentage points on XCOPA, 5.3 points on XNLI, and 1.9 points on MGSM.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.