GENESIS: 설명 가능한 인과 관계 추론을 향하여
GENESIS: Towards Explainable Causal Discovery
관측 데이터로부터의 인과 관계 추론(CD)은 두 가지 근본적인 어려움에 직면합니다. 첫째, 순수한 통계적 방법은 종종 낮은 샘플 환경에서 구조적 모호성을 해결할 수 있는 능력이 부족합니다. 둘째, LLM을 활용한 하이브리드 접근 방식이 의미론적 추론을 통해 구조 복원을 개선하지만, 이러한 추론이 개별적인 연결 결정에 미치는 영향은 여전히 불분명합니다. 결과적으로, 기존의 하이브리드 방법은 기본적인 요구 사항을 충족하지 못합니다. 즉, 학습된 방향성 비순환 그래프(DAG)에서 특정 연결이 포함되거나 제외되는 이유를 설명해야 한다는 것입니다. 이는 실제 응용 분야에서 매우 중요합니다. 왜냐하면 진정한 DAG가 존재하지 않고 모든 구조적 결정은 독립적으로 정당화되어야 하기 때문입니다. 우리는 이 요구 사항을 '결정 추적 가능성'으로 공식화하며, 추론된 모든 연결이 감사 가능한 통계적 증거, 마르코프 담요 일관성 또는 명시적인 도메인 지식을 통해 뒷받침되도록 합니다. 우리는 설명 가능한 하이브리드 CD 프레임워크인 GENESIS를 제안합니다. GENESIS는 그래프 구축을 해석 가능한 의사 결정 단계로 분해합니다. GENESIS는 먼저 체인, 갈림길 및 충돌 모티프와 같은 세 노드 구조적 모티프를 식별하고 점수를 매겨 투명한 구조적 사전 지식을 설정한 다음, 이러한 사전 지식을 관측 데이터 증거와 통합하여 그래프를 점진적으로 개선하며, 통계적 증거가 불충분할 때만 도메인 지식을 활용합니다. 설계상, 모든 연결 결정은 감사 가능한 출처의 증거를 통해 이루어집니다. 실험 결과는 GENESIS가 모든 설정에서 100%의 결정 추적 가능성을 달성하여 인과 관계 추론에서 설명 가능성을 최우선 목표로 삼았음을 보여줍니다. 이러한 추가적인 요구 사항에도 불구하고, GENESIS는 다양한 샘플 환경에서 대부분의 표준 데이터 세트에 대해 구조 해밍 거리(SHD) 측면에서 순수한 통계적 CD 방법보다 일관되게 우수한 성능을 보이며, 최첨단 LLM 지원 접근 방식과 유사한 성능을 달성합니다.
Causal Discovery (CD) from observational data faces two fundamental challenges. First, purely statistical methods often lack the power to resolve structural ambiguities in low-sample regimes. Second, although LLM-assisted hybrid approaches improve structure recovery through semantic reasoning, the influence of that reasoning on individual edge decisions remains largely opaque. Consequently, existing hybrid methods fail to satisfy a fundamental requirement: explaining why a particular edge is included or excluded in the learned directed acyclic graph (DAG). This is critical in real-world applications, where no ground-truth DAG exists and every structural decision must be independently justified. We formalize this requirement as decision traceability, requiring every inferred edge to be supported by auditable statistical evidence, Markov Blanket consistency, or explicit domain reasoning. We propose GENESIS, an explainable hybrid CD framework that decomposes graph construction into interpretable decision points. GENESIS first identifies and scores three-node structural motifs, including chains, forks, and colliders, to establish transparent structural priors, then progressively refines the graph by integrating these priors with observational evidence, invoking domain knowledge only when statistical evidence is insufficient. By design, every edge decision is resolved through an auditable source of evidence. Experiments show that GENESIS achieves 100% decision traceability across all settings, establishing explainability as a first-class objective in causal discovery. Despite this additional requirement, GENESIS consistently outperforms purely statistical CD methods on the majority of benchmark datasets across all sample regimes in terms of Structural Hamming Distance (SHD), while achieving performance comparable to state-of-the-art LLM-assisted approaches.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.