HLER: 다중 에이전트 파이프라인을 활용한 인간 협력 경제 연구를 통한 실증적 발견
HLER: Human-in-the-Loop Economic Research via Multi-Agent Pipelines for Empirical Discovery
대규모 언어 모델(LLM)은 과학 연구 워크플로우를 자동화하는 에이전트 기반 시스템을 가능하게 했습니다. 기존의 대부분 접근 방식은 완전 자율적인 발견에 초점을 맞추는데, 여기서 AI 시스템은 연구 아이디어를 생성하고, 분석을 수행하며, 최소한의 인간 개입으로 논문을 작성합니다. 그러나 경제학 및 사회과학 분야의 실증 연구는 추가적인 제약을 갖습니다. 연구 질문은 사용 가능한 데이터 세트에 기반해야 하며, 식별 전략은 신중한 설계가 필요하며, 인간의 판단은 경제적 중요성을 평가하는 데 필수적입니다. 본 논문에서는 인간 협력 경제 연구(HLER: Human-in-the-Loop Economic Research)라는 다중 에이전트 아키텍처를 소개합니다. 이 아키텍처는 실증 연구 자동화를 지원하는 동시에 중요한 인간의 감독을 유지합니다. 이 시스템은 데이터 감사, 데이터 프로파일링, 가설 생성, 계량 경제 분석, 논문 작성 및 자동 검토를 위한 특수 에이전트를 조정합니다. 핵심 설계 원칙은 데이터 세트 기반 가설 생성으로, 후보 연구 질문은 데이터 세트 구조, 변수 가용성 및 분포 진단에 의해 제한되어 실행 불가능하거나 허구적인 가설을 줄입니다. HLER은 또한 두 가지 루프 아키텍처를 구현합니다. 첫 번째는 실행 가능한 가설을 검토하고 선택하는 질문 품질 루프이고, 두 번째는 자동 검토가 재분석 및 논문 수정 작업을 트리거하는 연구 수정 루프입니다. 인간의 의사 결정 지점은 주요 단계에 내장되어 있어 연구자가 자동화된 파이프라인을 안내할 수 있습니다. 세 가지 실증 데이터 세트에 대한 실험 결과, 데이터 세트 기반 가설 생성은 87%의 경우에 실행 가능한 연구 질문을 생성합니다(제약 없는 생성의 경우 41%에 비해). 또한 평균 API 비용이 0.8달러에서 1.5달러로, 완전한 실증 논문을 생성할 수 있습니다. 이러한 결과는 인간-AI 협업 파이프라인이 확장 가능한 실증 연구를 위한 실용적인 경로를 제공할 수 있음을 시사합니다.
Large language models (LLMs) have enabled agent-based systems that aim to automate scientific research workflows. Most existing approaches focus on fully autonomous discovery, where AI systems generate research ideas, conduct analyses, and produce manuscripts with minimal human involvement. However, empirical research in economics and the social sciences poses additional constraints: research questions must be grounded in available datasets, identification strategies require careful design, and human judgment remains essential for evaluating economic significance. We introduce HLER (Human-in-the-Loop Economic Research), a multi-agent architecture that supports empirical research automation while preserving critical human oversight. The system orchestrates specialized agents for data auditing, data profiling, hypothesis generation, econometric analysis, manuscript drafting, and automated review. A key design principle is dataset-aware hypothesis generation, where candidate research questions are constrained by dataset structure, variable availability, and distributional diagnostics, reducing infeasible or hallucinated hypotheses. HLER further implements a two-loop architecture: a question quality loop that screens and selects feasible hypotheses, and a research revision loop where automated review triggers re-analysis and manuscript revision. Human decision gates are embedded at key stages, allowing researchers to guide the automated pipeline. Experiments on three empirical datasets show that dataset-aware hypothesis generation produces feasible research questions in 87% of cases (versus 41% under unconstrained generation), while complete empirical manuscripts can be produced at an average API cost of $0.8-$1.5 per run. These results suggest that Human-AI collaborative pipelines may provide a practical path toward scalable empirical research.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.