에이전트 기반 가설 수립 및 실험을 통한 사회과학 연구 가속화
Accelerating Social Science Research via Agentic Hypothesization and Experimentation
데이터 기반 사회과학 연구는 관찰, 가설 생성, 실험적 검증의 반복적인 주기에 의존하기 때문에 본질적으로 진행 속도가 느립니다. 최근의 데이터 기반 방법론들이 이 과정의 일부를 가속화할 수 있는 가능성을 보여주었지만, 종단간(end-to-end) 과학적 발견을 지원하는 데에는 대체로 미흡했습니다. 이러한 문제를 해결하기 위해, 우리는 생성자(Generator)가 후보 가설을 제안하고 실험자(Experimenter)가 이를 경험적으로 평가하는, 베이지안 최적화에서 영감을 받은 2단계 탐색을 통해 종단간 발견을 구현하는 에이전트 프레임워크인 EXPERIGEN을 소개합니다. 여러 도메인에 걸쳐 EXPERIGEN은 기존 접근 방식보다 예측력이 7~17% 더 높고 통계적으로 유의미한 가설을 2~4배 더 많이 일관되게 발견했으며, 멀티모달 및 관계형 데이터셋을 포함한 복잡한 데이터 환경으로도 자연스럽게 확장됩니다. 실질적인 과학적 진보를 이끌어내기 위해서는 통계적 성능을 넘어 가설이 참신하고, 경험적 근거가 있으며, 실행 가능해야 합니다. 이러한 특성을 평가하기 위해 우리는 기계가 생성한 가설에 대해 상급 교수진의 피드백을 수집하여 전문가 검토를 수행했습니다. 검토된 25개의 가설 중 88%가 중간 수준 이상으로 참신하다고 평가받았고, 70%는 영향력이 있으며 연구할 가치가 있는 것으로 판단되었으며, 대부분은 상급 대학원생 수준의 연구에 준하는 엄밀성을 보여주었습니다. 마지막으로, 궁극적인 검증을 위해서는 실제 세계의 증거가 필요하다는 점을 인식하여, LLM이 생성한 가설에 대해 최초로 A/B 테스트를 수행했으며, 그 결과 p값이 1e-6 미만이고 344%의 큰 효과 크기를 보이는 통계적으로 유의미한 결과를 확인했습니다.
Data-driven social science research is inherently slow, relying on iterative cycles of observation, hypothesis generation, and experimental validation. While recent data-driven methods promise to accelerate parts of this process, they largely fail to support end-to-end scientific discovery. To address this gap, we introduce EXPERIGEN, an agentic framework that operationalizes end-to-end discovery through a Bayesian optimization inspired two-phase search, in which a Generator proposes candidate hypotheses and an Experimenter evaluates them empirically. Across multiple domains, EXPERIGEN consistently discovers 2-4x more statistically significant hypotheses that are 7-17 percent more predictive than prior approaches, and naturally extends to complex data regimes including multimodal and relational datasets. Beyond statistical performance, hypotheses must be novel, empirically grounded, and actionable to drive real scientific progress. To evaluate these qualities, we conduct an expert review of machine-generated hypotheses, collecting feedback from senior faculty. Among 25 reviewed hypotheses, 88 percent were rated moderately or strongly novel, 70 percent were deemed impactful and worth pursuing, and most demonstrated rigor comparable to senior graduate-level research. Finally, recognizing that ultimate validation requires real-world evidence, we conduct the first A/B test of LLM-generated hypotheses, observing statistically significant results with p less than 1e-6 and a large effect size of 344 percent.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.