2605.29411v1 May 28, 2026 cs.LG

테이블 예측을 위한 마르코프 경계의 장점, 단점 및 문제점

The Good, the Bad, and the Ugly of Markov Boundary for Tabular Prediction

Shu Wan
Shu Wan
Citations: 33
h-index: 2
Huan Liu
Huan Liu
Citations: 23
h-index: 2
Abhinav Gorantla
Abhinav Gorantla
Citations: 6
h-index: 2
K. Candan
K. Candan
Citations: 10
h-index: 2

표준 그래프 가정 하에서, 대상 변수의 마르코프 경계는 다른 모든 특징을 불필요하게 만드는 최소한의 특징 집합입니다. 이 경계를 관찰하면 대상은 나머지 테이블과 조건적으로 독립적입니다. 이는 테이블 예측에 매력적인 요소이며, 모델이 실제로 필요로 하는 열을 정확히 지칭합니다. 그러나 현대 회귀 모델은 여전히 전체 특징 집합으로 훈련됩니다. 본 연구에서는 SCM3K라는 3,450개의 작업으로 구성된 합성 SCM 벤치마크에서 마르코프 경계가 예측에 실제로 유용한지 조사했습니다. 이 벤치마크는 특징 개수가 40개에서 1000개까지 다양하며, 여섯 가지 SCM 패밀리로 구성되어 있습니다. 또한, 여섯 개의 회귀 모델을 사용하여 성능을 평가했습니다. 결과는 이론이 제시하는 것보다 더 복잡합니다. 회귀 모델을 오라클 경계에 제한하면 예측 성능이 크게 향상되는 경우가 많으며, 특징 공간의 크기와 희소성이 증가할수록 이러한 개선 효과는 더욱 커집니다. 그러나 인과 추론을 통해 경계를 복구하고 복구된 마스크를 사용하여 훈련하는 일반적인 파이프라인은 원하는 결과를 얻지 못합니다. 기존 추정기는 경계가 가장 도움이 되는 시점에 도달하기 전에 계산 예산을 모두 소모하며, 실행되더라도 전체 특징 집합보다 성능이 뛰어난 경우는 드뭅니다. 이러한 현상은 다음 세 가지 원인으로 인해 발생합니다. 첫째, 인과 추론은 구조 복구에 최적화되어 예측 성능에는 직접적으로 기여하지 않습니다. 둘째, 오탐(false positive)과 미탐(false negative)은 예측 비용 측면에서 비대칭적인 영향을 미칩니다. 셋째, 정확한 경계는 모든 특징보다 더 나은 성능을 보이는 많은 특징 집합 중 하나일 뿐입니다. 이러한 사실들을 바탕으로, 예측에 최적화된 특징 선택 및 인과 구조를 활용하는 테이블 모델 개발을 위한 시사점을 논의합니다.

Original Abstract

Under standard graphical assumptions, the Markov boundary of a target variable is the smallest set of features that renders every other feature redundant. Once the boundary is observed, the target is conditionally independent of the rest of the table. This is a tempting object for tabular prediction, since it names exactly the columns a model should need. Yet modern regressors are still trained on the full feature set. We ask whether the Markov boundary is genuinely useful for prediction on SCM3K, a 3,450-task synthetic SCM benchmark with feature counts from 40 to 1000 and six SCM families, evaluated with six regressors. The answer is more nuanced than the theory suggests. Restricting a regressor to the oracle boundary often improves prediction substantially, and the improvement grows as the feature space becomes larger and sparser. But the natural pipeline of recovering the boundary with causal discovery and training on the recovered mask does not deliver. Existing estimators exhaust the compute budget before reaching the regime where the boundary helps most, and even where they run they rarely beat the full feature set. We trace this to three causes. Discovery optimizes structural recovery rather than prediction. False negatives and false positives carry sharply asymmetric predictive cost. The exact boundary is only one of many feature sets that beat all features. We then develop what these facts imply for prediction-aligned feature selection and for tabular models that learn to use causal structure.

0 Citations
0 Influential
1 Altmetric
5.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!