AVA로부터 배우는 교훈: 정책 및 개발 연구를 위한 선별되고 신뢰할 수 있는 생성형 AI의 초기 사례
Learning from AVA: Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development Research
범용 LLM은 개발 및 정책 전문가에게 오정보 위험을 초래하며, 검증 가능한 결과물을 생성하기 위한 지적 겸손이 부족합니다. 본 연구에서는 4,000개 이상의 세계은행 보고서로 구성된 선별된 라이브러리를 기반으로 다국어 기능을 갖춘 생성형 AI 플랫폼인 AVA (AI + Verified Analysis)를 소개합니다. AVA의 멀티 에이전트 파이프라인은 사용자가 질의를 하고 증거 기반의 종합적인 결과를 얻을 수 있도록 합니다. AVA는 인용 가능성 확인(주장에 대한 출처 추적)과 근거 있는 거부(근거 없는 질의에 대한 거부 및 정당화와 리디렉션)라는 두 가지 메커니즘을 통해 지적 겸손을 구현합니다. 우리는 116개 국가의 다양한 조직 및 역할을 가진 2,200명 이상의 사용자를 대상으로 로그 분석, 설문 조사 및 20건의 인터뷰를 통해 실제 환경에서의 평가를 수행했습니다. Difference-in-Differences 분석 결과, 지속적인 사용은 주당 2.4~3.9시간의 시간을 절약하는 것으로 나타났습니다. 또한, 참가자들은 AVA를 전문적인 "증거 엔진"으로 사용했으며, 근거 있는 거부는 범위의 경계를 명확히 하고, 기관적 출처 및 페이지 단위 인용을 통해 신뢰도를 높였습니다. 본 연구는 전문 AI 설계를 위한 지침을 제시하고, "생태계 인식"을 갖춘 겸손한 AI에 대한 비전을 제시합니다.
General-purpose LLMs pose misinformation risks for development and policy experts, lacking epistemic humility for verifiable outputs. We present AVA (AI + Verified Analysis), a GenAI platform built on a curated library of over 4,000 World Bank Reports with multilingual capabilities. AVA's multi-agent pipeline enables users to query and receive evidence-based syntheses. It operationalizes epistemic humility through two mechanisms: citation verifiability (tracing claims to sources) and reasoned abstention (declining unsupported queries with justification and redirection). We conducted an in-the-wild evaluation with over 2,200 individuals from heterogeneous organisations and roles in 116 countries, via log analysis, surveys, and 20 interviews. Difference-in-Differences estimates associate sustained engagement with 2.4-3.9 hours saved weekly. Qualitatively, participants used AVA as a specialized "evidence engine"; reasoned abstention clarified scope boundaries, and trust was calibrated through institutional provenance and page-anchored citations. We contribute design guidelines for specialized AI and articulate a vision for "ecosystem-aware" Humble AI.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.