언어 모델 내 숨겨진 API: 분기된 미래로부터 재사용 가능한 인과적 인터페이스 발견
Hidden APIs in Language Models: Discovering Reusable Causal Interfaces from Forked Futures
동일한 언어 모델 응답은 다양한 미래 연산을 지원하는 숨겨진 상태에서 발생할 수 있으므로, 현재 응답 기반의 분석 방법으로는 재사용 가능한 내부 인터페이스를 확립하기 어렵습니다. 본 연구에서는 '분기된 미래(forked futures)'라는 개념을 도입합니다. 분기된 미래는 접두사 상태가 형성된 후에 미래 연산을 샘플링하며, 해당 연산에 의해 유도되는 응답 분포를 통해 상태들을 비교합니다. 이를 통해 연구자가 직접 지정한 잠재적 레이블 없이 숨겨진 상태에 대한 경험적인 인과적 지표를 얻을 수 있습니다. 공유형(Shared), 지역형(Local), 혼합형(Mixture), 분산형(Distributed) 인터페이스는 미래의 특징 충실도와 일관된 용량 제약 조건 하에서, 사전 예측 인과적 설명 길이(prequential causal description length)를 기준으로 경쟁합니다. 두 가지 모델 평가 결과, 공유형 인터페이스가 가장 낮은 설명 길이를 보였습니다. 구체적으로 Qwen2.5-1.5B 모델에서 0.216 nats, Llama-3-8B 모델에서 0.294 nats의 성능 향상을 보였으며, 동시에 미래 특징 왜곡이 매우 작게 유지되었습니다. 다섯 가지 핵심 구조에 대한 분석 결과, 공유형 인터페이스가 가장 높은 정확성, 지역성, 복사 보존성 및 통합 프로필을 나타냈습니다. API와 일치하는 경로는 목표 효과의 0.749를 담당하는 반면, 일치된 무작위 경로의 경우 0.150에 불과했습니다. 실제 모델-생명체 분류 테스트에서는 16개 아키텍처 중 14개가 정확하게 복원되었으며, 12개의 비-공유형 생명체 중에서 한 가지가 공유형으로 잘못 분류되는 오류가 관찰되었습니다. 이러한 결과는 테스트된 연산 집합 내에서 경제적인 재사용 가능한 인과적 인터페이스의 존재 가능성을 시사합니다. 하지만 본 연구 결과는 검토된 아키텍처, 개입 및 보류된 미래에 대한 명시적인 조건부 주장을 포함하고 있습니다.
Identical language-model answers can arise from hidden states that support different future computations, so current-answer probes do not establish a reusable internal interface. We introduce forked futures: future operations are sampled only after a prefix state has formed, and states are compared through the response distributions induced by those operations. This yields an empirical causal quotient over hidden states without requiring researcher-specified latent labels. Shared, Local, Mixture, and Distributed interfaces then compete under prequential causal description length subject to future-signature fidelity and matched capacity constraints. In the two detailed model evaluations, Shared has the lowest held-out description length, with gains of 0.216 nats on Qwen2.5-1.5B and 0.294 nats on Llama-3-8B, while maintaining tightly clustered mean future-signature distortion; a five-backbone sweep preserves the positive direction of Sharedness Gain. The figure-aligned transplantation analysis gives Shared the strongest joint target-correctness, locality, copy-preservation, and composite profile, and API-aligned paths mediate 0.749 of the target effect versus 0.150 for matched null paths. In the blind four-class model-organism test, 14/16 architectures are recovered, with one observed non-Shared to Shared error among 12 non-Shared organisms. These results support an economical reusable causal interface within the tested operation banks, while keeping the claim explicitly conditional on the candidate architectures, interventions, and held-out futures.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.