CausaLab: AI 연구자를 위한 확장 가능한 대화형 인과 관계 추론 환경
CausaLab: A Scalable Environment for Interactive Causal Discovery Toward AI Scientists
본 논문에서는 LLM 에이전트가 참여하는 대화형 인과 관계 추론을 평가하기 위한 확장 가능한 환경인 CausaLab을 소개합니다. 기존의 평가 방식과는 달리, CausaLab은 에이전트가 인과적 증거를 활용하여 문제를 해결할 수 있는지 여부와 함께, 에이전트의 답변이 기본 인과 메커니즘에 대한 정확한 가설에 의해 뒷받침되는지 여부를 평가합니다. 각 실험 단계에서 에이전트는 합성 실험실 환경에 배치되며, 사전 측정 기록을 받고, 조작 결정을 내리고, 동일한 메커니즘에 따라 작동하는 숨겨진 반응기 결정의 공진 주파수를 예측합니다. 데이터 생성 프로세스는 무작위로 샘플링된 구조적 인과 모델(SCM)이며, 성공적인 결과를 얻으려면 사전 지식을 회상하는 것이 아니라 인과 그래프와 구조 방정식을 모두 복원해야 합니다. CausaLab은 또한 에이전트의 진화하는 SCM 가설을 기록하는 도메인 특화 언어를 포함하여, 추론 과정을 검사하고 실제 값과 비교할 수 있도록 지원합니다. 실험 결과, 예측 정확도와 메커니즘 복구 사이에 지속적인 격차가 존재함을 확인했습니다. 순수 관찰 데이터만 사용한 6개 노드 환경에서 GPT-5.2-high는 92%의 작업 정확도를 달성했지만, 모든 연결에 대한 $F_1$ 점수는 0.471에 불과했습니다. 이러한 관찰은 다양한 상호 작용 전략을 탐구하도록 유도했습니다. 관찰 데이터와 조작 결정을 혼합한 전략은 구조적 정확도를 향상시켰습니다. 혼합된 6개 노드 환경에서 GPT-5.2-high는 작업 정확도와 모든 연결에 대한 $F_1$ 점수 모두에서 80%를 달성했습니다. 하지만 강력한 에이전트조차도 유용한 조작을 설계하는 데 어려움을 겪으며, 순수한 조작 전략은 작업 정확도와 모든 연결에 대한 $F_1$ 점수 측면에서 모두 낮은 성능을 보였습니다. 본 연구에서는 에이전트의 주요 약점 중 하나를 조기 종료로 지목하고, 모델에게 자신의 가설과 과거 데이터 간의 일관성을 검증하도록 요청하면 이 문제를 완화하는 데 도움이 될 수 있음을 보여주었습니다. 따라서 CausaLab은 예측 성공과 인과적 이해 사이의 연관성을 분리하며, 현재 LLM 에이전트가 실험적인 인과 추론 도구로서 갖는 한계를 드러냅니다.
We introduce CausaLab, a scalable environment for evaluating interactive causal discovery by LLM agents. Unlike prior evaluations, CausaLab evaluates both whether an agent can solve a problem using causal evidence and whether its answer is supported by a correct hypothesis about the underlying causal mechanism. Each episode places an agent in a synthetic laboratory: it receives prior measurement records, intervenes on a manipulator crystal, and predicts the resonance frequency of a held-out reactor crystal governed by the same mechanism. The hidden data-generating process is a randomly sampled structural causal model (SCM), so success requires recovering both a causal graph and structural equations rather than recalling prior knowledge. CausaLab also includes a domain-specific language that records the agent's evolving SCM hypothesis, making trajectories inspectable and comparable with ground truth. Experiments show a persistent gap between prediction and mechanism recovery: in the purely observational 6-node setting, GPT-5.2-high reaches 92% task accuracy but only 0.471 all-edge $F_1$. This observation further motivates our exploration of different interaction strategies: Mixed observation--intervention strategies improve structural fidelity: in the mixed 6-node setting, GPT-5.2-high achieves 80% on both task accuracy and all-edge $F_1$. Yet even strong agents struggle to design informative interventions, as pure intervention strategies perform poorly on both task accuracy and all-edge $F_1$. We identify premature stopping as a major weakness of agents, and show that asking the model to verify the consistency between its hypothesis and past data can help mitigate this issue. CausaLab therefore separates predictive success from causal understanding and exposes current LLM agents' limits as experimental causal reasoners.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.