공학 시뮬레이션 파이프라인에서의 순차적 제약 조건 수정 작업을 위한 물리 기반 인과 관계 마르코프 결정 프로세스(CMDP)
Physics-Informed Causal MDPs for Sequential Constraint Repair in Engineering Simulation Pipelines
제약 조건이 있는 대규모 이진 상태 공간을 가진 마르코프 결정 프로세스(MDP)에서의 오프라인 학습은 근본적인 어려움을 겪습니다. 인과 관계를 파악하기 위해서는 구조적 가정이 필요하지만, 샘플 효율적인 정책 학습을 위해서는 상태 공간의 축소가 필요하기 때문입니다. 본 연구에서는 라이프사이클 순서 가정(LOA) 하에서 제약 조건 간의 의존성이 계층적 방향성 비순환 그래프(DAG)를 형성하는 CMDP 프레임워크인 PI-CMDP를 제안합니다. 우리는 '식별-압축-추정' 파이프라인을 제안합니다. (i) 식별 단계: LOA는 층 간 쌍에 대한 인과 관계 에ッジ 가중치를 백도어 식별을 통해 파악할 수 있으며, LOA가 위반될 경우에도 엄격한 부분 식별 경계를 제공합니다. (ii) 압축 단계: 마르코프 추상화는 층 우선 순위 규칙성과 교환성을 통해 상태 공간의 크기를 2^(WL)에서 (W+1)^L로 줄입니다. (iii) 추정 단계: 물리 기반 이중 강건 추정기는 물리적 사전 지식이 학습된 모델보다 우수한 경우에도 편향되지 않으며, 분산 상수를 줄입니다. PI-CMDP는 공학 시뮬레이션 파이프라인에서의 제약 조건 수정 작업에 적용되었습니다. TPS 벤치마크(4,206 에피소드)에서 PI-CMDP는 300개의 학습 에피소드만으로 76.2%의 수정 성공률을 달성했으며, 이는 가장 강력한 기준 모델(70.8%)보다 5.4%p 더 높은 수치입니다. 전체 데이터 환경에서는 2.8%p(83.4% vs. 80.6%)의 성능 향상을 보였으며, 또한 연쇄적인 오류 발생률을 크게 줄였습니다. 이러한 모든 개선 사항은 5개의 독립적인 시드(paired t-test p < 0.02)에 걸쳐 일관되게 나타났습니다.
Off-policy learning in constrained MDPs with large binary state spaces faces a fundamental tension: causal identification of transition dynamics requires structural assumptions, while sample-efficient policy learning requires state-space compression. We introduce PI-CMDP, a framework for CMDPs whose constraint dependencies form a layered DAG under a Lifecycle Ordering Assumption (LOA). We propose an Identify-Compress-Estimate pipeline: (i) Identify: LOA enables backdoor identification of causal edge weights for cross-layer pairs, with formal partial-identification bounds when LOA is violated; (ii) Compress: a Markov abstraction compresses state cardinality from 2^(WL) to (W+1)^L under layer-priority regularity and exchangeability; and (iii) Estimate: a physics-guided doubly-robust estimator remains unbiased and reduces the variance constant when the physics prior outperforms a learned model. We instantiate PI-CMDP on constraint repair in engineering simulation pipelines. On the TPS benchmark (4,206 episodes), PI-CMDP achieves 76.2% repair success rate with only 300 training episodes versus 70.8% for the strongest baseline (+5.4 pp), narrowing to +2.8 pp (83.4% vs. 80.6%) in the full-data regime, while substantially reducing cascade failure rates. All improvements are consistent across 5 independent seeds (paired t-test p < 0.02).
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.