2604.17843v1 Apr 20, 2026 cs.HC

AVA로부터 배우는 교훈: 정책 및 개발 연구를 위한 선별되고 신뢰할 수 있는 생성형 AI의 초기 사례

Learning from AVA: Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development Research

Nimisha Karnatak
Nimisha Karnatak
Citations: 19
h-index: 3
Mohamad Chatila
Mohamad Chatila
Citations: 5
h-index: 1
Daniel Alejandro Pinzón Hernández
Daniel Alejandro Pinzón Hernández
Citations: 179
h-index: 1
R. Yazdanfar
R. Yazdanfar
Citations: 76
h-index: 1
Michelle Dugas
Michelle Dugas
Citations: 2
h-index: 1
Renos Vakis
Renos Vakis
Citations: 3,521
h-index: 27

범용 LLM은 개발 및 정책 전문가에게 오정보 위험을 초래하며, 검증 가능한 결과물을 생성하기 위한 지적 겸손이 부족합니다. 본 연구에서는 4,000개 이상의 세계은행 보고서로 구성된 선별된 라이브러리를 기반으로 다국어 기능을 갖춘 생성형 AI 플랫폼인 AVA (AI + Verified Analysis)를 소개합니다. AVA의 멀티 에이전트 파이프라인은 사용자가 질의를 하고 증거 기반의 종합적인 결과를 얻을 수 있도록 합니다. AVA는 인용 가능성 확인(주장에 대한 출처 추적)과 근거 있는 거부(근거 없는 질의에 대한 거부 및 정당화와 리디렉션)라는 두 가지 메커니즘을 통해 지적 겸손을 구현합니다. 우리는 116개 국가의 다양한 조직 및 역할을 가진 2,200명 이상의 사용자를 대상으로 로그 분석, 설문 조사 및 20건의 인터뷰를 통해 실제 환경에서의 평가를 수행했습니다. Difference-in-Differences 분석 결과, 지속적인 사용은 주당 2.4~3.9시간의 시간을 절약하는 것으로 나타났습니다. 또한, 참가자들은 AVA를 전문적인 "증거 엔진"으로 사용했으며, 근거 있는 거부는 범위의 경계를 명확히 하고, 기관적 출처 및 페이지 단위 인용을 통해 신뢰도를 높였습니다. 본 연구는 전문 AI 설계를 위한 지침을 제시하고, "생태계 인식"을 갖춘 겸손한 AI에 대한 비전을 제시합니다.

Original Abstract

General-purpose LLMs pose misinformation risks for development and policy experts, lacking epistemic humility for verifiable outputs. We present AVA (AI + Verified Analysis), a GenAI platform built on a curated library of over 4,000 World Bank Reports with multilingual capabilities. AVA's multi-agent pipeline enables users to query and receive evidence-based syntheses. It operationalizes epistemic humility through two mechanisms: citation verifiability (tracing claims to sources) and reasoned abstention (declining unsupported queries with justification and redirection). We conducted an in-the-wild evaluation with over 2,200 individuals from heterogeneous organisations and roles in 116 countries, via log analysis, surveys, and 20 interviews. Difference-in-Differences estimates associate sustained engagement with 2.4-3.9 hours saved weekly. Qualitatively, participants used AVA as a specialized "evidence engine"; reasoned abstention clarified scope boundaries, and trust was calibrated through institutional provenance and page-anchored citations. We contribute design guidelines for specialized AI and articulate a vision for "ecosystem-aware" Humble AI.

2 Citations
0 Influential
13.5 Altmetric
69.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!