2606.17516v1 Jun 16, 2026 cs.LG

FoundCause: 관찰 데이터로부터 잠재적 교란 변수를 활용한 인과 관계 추론

FoundCause: Causal Discovery with Latent Confounders from Observational Data

S. Kasiviswanathan
S. Kasiviswanathan
Citations: 5,935
h-index: 28
K. Balasubramanian
K. Balasubramanian
Citations: 18
h-index: 3
Patrick Blobaum
Patrick Blobaum
Citations: 22
h-index: 2

관찰 데이터에서 인과 관계를 추론하는 것은 개입 없이 방향성 구조와 잠재적인 교란 요인을 복구해야 하므로 여전히 어려운 과제입니다. 본 논문에서는 FoundCause라는 새로운 모델을 제안합니다. FoundCause는 합성 데이터를 기반으로 학습된, 모든 과정을 한 번의 순방향 연산으로 데이터셋을 직접 인과 그래프로 매핑하는 효율적인 인과 관계 추론 모델입니다. FoundCause는 다양한 시뮬레이션된 구조적 인과 모델(structural causal models)로부터 학습하여 개별 데이터셋을 넘어 일반화될 수 있는 전이 가능한 통계적 패턴을 파악합니다. 이 모델은 인과 관계 추론에 중요한 여러 가지 사전 지식(inductive biases)을 통합하고 있습니다. 특히, 변수와 샘플 간의 상호 의존성 및 각 변수의 분포를 동시에 모델링하기 위해 순열 불변 트랜스포머 인코더와 교차 변수 주의 메커니즘을 사용합니다. 고전적인 비대칭 측정법에서 파생된 쌍별 통계적 특징은 통계 조건부 주의(statistics-conditioned attention)를 통해 주입되어 모델이 알려진 인과적 신호를 향하도록 안내합니다. 팩토리화된 디코더는 간선의 존재 여부와 방향을 분리하고, 삼각 개선 모듈(triangular refinement module)은 연쇄 및 충돌자(collider)와 같은 고차원 인과적 패턴에 대한 추론을 가능하게 합니다. 또한, 학습 가능한 잠재 토큰 기반의 전용 교란 변수 모듈은 숨겨진 공통 원인을 명시적으로 모델링하며, 마스킹된 입력 표현 방식을 통해 누락된 데이터도 처리합니다. 현재까지 알려진 바로는, FoundCause는 잠재적인 교란 요인을 명시적으로 모델링하는 최초의 효율적인 인과 관계 추론 방법입니다. FoundCause는 15개의 실제 데이터셋에 대해 11가지 기존 비효율적인 방법(예: PC, GES, NOTEARS 스타일 최적화) 및 4가지 다른 효율적인 방법보다 $F_1$ 점수에서 +9.6%, AUROC에서 +1.2% 향상되었으며, 가장 성능이 좋은 비효율적인 방법에 비해 구조 해밍 거리(structural Hamming distance)가 18.9% 감소했습니다. 또한 FoundCause는 단일 순방향 연산으로 추론을 수행합니다.

Original Abstract

Causal discovery from observational data remains challenging due to the need to recover directed structure and latent confounding without interventions. We propose FoundCause, an amortized causal discovery model trained entirely on synthetic data that maps datasets directly to causal graphs in a single forward pass. By learning from large collections of simulated structural causal models, FoundCause captures transferable statistical patterns that generalize beyond individual datasets. The architecture incorporates several key inductive biases for causal discovery. It uses a permutation-invariant transformer encoder with alternating attention over samples and variables to jointly model cross-variable dependence and per-variable distributions. Pairwise statistical features derived from classical asymmetry measures are injected through statistics-conditioned attention, guiding the model toward known causal signals. A factorized decoder separates edge existence from direction, while a triangular refinement module enables reasoning over higher-order causal motifs such as chains and colliders. In addition, a dedicated confounder module based on learnable latent tokens explicitly models hidden common causes, and the model explicitly handles missing data via its masked input representation. To our knowledge, FoundCause is the first amortized causal discovery approach to explicitly model latent confounding. FoundCause outperforms 11 classical non-amortized methods (e.g., PC, GES, NOTEARS-style optimization) and 4 amortized causal discovery methods on 15 real-world datasets, achieving +9.6% improvement in $F_1$, +1.2% in AUROC, and an 18.9% reduction in structural Hamming distance relative to the strongest non-amortized methods, while performing inference in a single forward pass.

1 Citations
1 Influential
14 Altmetric
73.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!