SynthSAEBench: 확장 가능한 현실적인 합성 데이터를 활용한 희소 자동 인코더 평가
SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data
희소 자동 인코더(SAE)의 성능 향상을 위해서는 아키텍처 혁신을 정확하게 검증할 수 있는 벤치마크가 필요합니다. 하지만 현재 LLM(Large Language Models, 거대 언어 모델)에 사용되는 SAE 벤치마크는 종종 노이즈가 많아 아키텍처 개선 사항을 구별하기 어렵고, 현재의 합성 데이터 실험은 규모가 작고 현실적이지 않아 의미 있는 비교를 제공하지 못합니다. 본 연구에서는 현실적인 특징(상관관계, 계층 구조, 중첩 등)을 가진 대규모 합성 데이터를 생성하는 도구인 SynthSAEBench와, SAE 아키텍처의 직접적인 비교를 가능하게 하는 표준화된 벤치마크 모델인 SynthSAEBench-16k를 소개합니다. 우리의 벤치마크는 재구성 성능과 잠재 공간 품질 지표 간의 불일치, SAE 프로빙 결과의 열악함, 그리고 L0 정밀도-재현율 균형과 같은 이전에 관찰된 LLM SAE 현상을 재현합니다. 또한, 우리의 벤치마크를 사용하여 새로운 실패 모드를 식별했습니다. Matching Pursuit SAE는 실제 특징을 학습하지 않고 중첩 노이즈를 활용하여 재구성 성능을 향상시키는데, 이는 더 표현력이 뛰어난 인코더가 쉽게 과적합될 수 있음을 시사합니다. SynthSAEBench는 실제 특징과 통제된 실험을 제공함으로써 LLM 벤치마크를 보완하고, 연구자들이 SAE의 실패 원인을 정확하게 진단하고 LLM으로 확장하기 전에 아키텍처 개선 사항을 검증할 수 있도록 지원합니다.
Improving Sparse Autoencoders (SAEs) requires benchmarks that can precisely validate architectural innovations. However, current SAE benchmarks on LLMs are often too noisy to differentiate architectural improvements, and current synthetic data experiments are too small-scale and unrealistic to provide meaningful comparisons. We introduce SynthSAEBench, a toolkit for generating large-scale synthetic data with realistic feature characteristics including correlation, hierarchy, and superposition, and a standardized benchmark model, SynthSAEBench-16k, enabling direct comparison of SAE architectures. Our benchmark reproduces several previously observed LLM SAE phenomena, including the disconnect between reconstruction and latent quality metrics, poor SAE probing results, and a precision-recall trade-off mediated by L0. We further use our benchmark to identify a new failure mode: Matching Pursuit SAEs exploit superposition noise to improve reconstruction without learning ground-truth features, suggesting that more expressive encoders can easily overfit. SynthSAEBench complements LLM benchmarks by providing ground-truth features and controlled ablations, enabling researchers to precisely diagnose SAE failure modes and validate architectural improvements before scaling to LLMs.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.