2606.13670v1 Jun 11, 2026 cs.AI

대규모 언어 모델을 활용한 사회 및 행동 과학 분야의 자동화된 재현성 평가

Automated reproducibility assessments in the social and behavioral sciences using large language models

Frauke Kreuter
Frauke Kreuter
Citations: 383
h-index: 11
S. Feuerriegel
S. Feuerriegel
Citations: 10,075
h-index: 48
Tobias Holtdirk
Tobias Holtdirk
Citations: 3
h-index: 1
Pietro Marcolongo
Pietro Marcolongo
Citations: 0
h-index: 0
Anna Schulten
Anna Schulten
Citations: 137
h-index: 7
Felix Henninger
Felix Henninger
Citations: 8,002
h-index: 14
Stefan Rose
Stefan Rose
Citations: 122
h-index: 4
Sarah Ball
Sarah Ball
Citations: 66
h-index: 4
Bolei Ma
Bolei Ma
Citations: 12
h-index: 2
Markus Weinmann
Markus Weinmann
Citations: 18
h-index: 1

사회 및 행동 과학 분야에서 재현성은 일반적으로 독립적인 연구자들이 원본 데이터를 재분석하여 발표된 결과가 재현 가능한지 평가하는 방식으로 이루어진다. 그러나 이러한 방식은 자원 집약적이며 확장하기 어렵다. 본 논문에서는 대규모 언어 모델(LLM)을 사용하여 재현성 평가를 자동화할 수 있음을 보여준다. 행동 및 사회 과학 분야에서 미리 정의된 가설을 가진 76편의 발표된 연구를 대상으로, LLM이 생성한 분석 결과와 원본 결과, 그리고 인간 전문가의 재분석 결과를 비교하였다. 7개의 연구에서는 LLM이 유효한 효과 크기 추정치를 산출하지 못했다. 나머지 연구에서, LLM 파이프라인은 Cohen's d 기준 +/-0.05의 허용 오차 내에서 원본 효과 크기를 41%의 연구에서 재현하는 데 성공했다. 또한, LLM 파이프라인은 96%의 경우에 원본 연구와 동일한 질적인 결론을 도출했으며, 여기서 결론은 재분석 결과가 원본 가설을 뒷받침하는지 여부를 나타낸다. 비교를 위해, 인간 전문가의 재분석은 34%의 연구에서 원본 효과 크기를 재현하고 74%의 경우에 동일한 질적인 결론을 도출했다. 이러한 결과를 종합적으로 고려할 때, LLM은 사회 및 행동 과학 분야의 실증적 결과에 대한 체계적인 검토를 위한 확장 가능한 도구로 사용될 수 있으며, 자동화된 재현성 평가의 기반을 제공한다.

Original Abstract

Reproducibility in the social and behavioral sciences is typically evaluated by independent researchers who reanalyze the original data to assess whether the published findings can be recovered. However, such approaches are resource-intensive and difficult to scale. Here, we show that large language models (LLMs) can automate reproducibility assessments. Using N=76 published studies with predefined claims from the behavioral and social sciences, we compare LLM-generated analysis with the original findings and human reanalysis. For 7 studies, the LLM could not produce a viable effect size estimate. For the remaining studies, our LLM pipeline recovered the original effect sizes in 41% of studies using a +/-0.05 tolerance in Cohen's d. Further, our LLM pipeline reached the same qualitative conclusion as the original study in 96% of cases, where conclusions indicate whether the reanalysis supports the original claim. For comparison, human reanalysts recovered the original effect sizes in 34% of studies and reached the same qualitative conclusion in 74% of cases. Together, these results show that LLMs can serve as a scalable tool for automated reproducibility assessment and provide a foundation for systematic auditing of empirical results in the social and behavioral sciences.

1 Citations
0 Influential
24 Altmetric
121.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!