트랜스포머를 활용한 베이지안 인-컨텍스트 실험: 평활도 적응형 효율적인 평균 치료 효과 추정
Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
평균 치료 효과(ATE)에 대한 적응형 실험은 유효한 추론과 통계적 효율성을 균형 있게 유지하는 랜덤 할당이 필요합니다. 이상적인 설계는 알려지지 않은 조건부 결과 분산을 기반으로 하는 변수 의존적인 네이먼 규칙을 따릅니다. 본 연구에서는 이러한 순차적인 분산 추정 및 할당 프로세스가 인-컨텍스트 학습을 통해 효율화될 수 있는지 조사합니다. 우리는 베이지안 인-컨텍스트 실험기를 소개합니다. 이는 베이지안 후류 네이먼 모델을 모방하도록 훈련된 트랜스포머 정책입니다. 이 모델은 실험 기록을 사용하여 잠재적 결과에 대한 비모수적 믿음을 업데이트하고, 후류 네이먼 치료 확률을 할당합니다. 이러한 설계는 이상적인 규칙으로 수렴하며, 효율적인 ATE 추론을 지원합니다. 트랜스포머는 어텐션 기반의 충분 통계량과 투영 경사 하강법을 사용하여 이 매핑을 구현하고, 가우시안 시리즈 사전 분포에 대한 베이지안 업데이트를 모방합니다. 알려지지 않은 결과의 평활도를 해결하기 위해, 우리는 평활도 지수를 사용하는 실험기를 결합하여 믹스처-오브-익스퍼트 트랜스포머를 사용합니다. 게이트는 평활도 클래스에 대한 계층적 후류 분포 역할을 하며, 이상적인 성능을 보이는 전문가에게 집중합니다. 트랜스포머 클래스의 복잡성을 제한함으로써, 우리는 이 효율화된 정책이 감독 사전 훈련을 사용하여 경험적 위험 최소화를 통해 학습될 수 있음을 증명합니다. 실험 결과는 정확한 모델 모방, 적응형 할당 및 기존 방법보다 향상된 ATE 정밀도를 확인했습니다.
Adaptive experiments for average treatment effects (ATE) require randomized allocations balancing valid inference with statistical efficiency. The oracle design is a covariate-dependent Neyman rule governed by unknown arm-conditional outcome variances. We investigate whether this sequential variance-estimation and allocation process can be amortized via in-context learning. We introduce Bayesian in-context experimenters: transformer policies trained to imitate a Bayesian posterior Neyman teacher. The teacher updates nonparametric beliefs over potential outcomes using experimental history to assign posterior Neyman treatment probabilities. This design converges to the oracle rule, supporting efficient ATE inference. Transformers constructively implement this mapping through attention-based sufficient statistics and projected gradient descent, imitating Bayesian updating for Gaussian-series priors. To address unknown outcome smoothness, we combine smoothness-indexed experimenters using a mixture-of-experts transformer. The gate acts as a hierarchical posterior over smoothness classes, concentrating on near-oracle experts. By bounding the complexity of the transformer class, we prove this amortized policy can be learned via empirical risk minimization using supervised pretraining. Experiments confirm accurate teacher imitation, adaptive allocation, and improved ATE precision over baselines.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.