2606.19782v1 Jun 18, 2026 cs.AI

AgentFinVQA: 감사 가능한 금융 차트 질의응답을 위한 배포 가능한 다중 에이전트 파이프라인

AgentFinVQA: A Deployable Multi-Agent Pipeline for Auditable Financial Chart QA

Aravind Narayanan
Aravind Narayanan
Citations: 27
h-index: 3
Shaina Raza
Shaina Raza
Citations: 13
h-index: 2

규제 환경에서 금융 차트 질의응답은 정확성뿐만 아니라 신뢰성을 요구합니다. 실무자는 어떤 답변을 믿어야 할지 판단해야 하며, 많은 기관은 고객 데이터를 외부 모델 제공업체로 전송할 수 없습니다. 그러나 기존의 차트 질의응답 에이전트는 주로 정확성에 초점을 맞추고 있으며 투명성이 부족하며, 대부분 독점적인 API 접근 방식을 사용합니다. 현재까지 알려진 바로는, 감사 가능성과 온프레미스 배포 가능성을 결합하면서도 정확도를 크게 저해하지 않는 시스템은 존재하지 않습니다. 본 논문에서는 AgentFinVQA를 소개합니다. 이는 각 질의를 계획, OCR (광학 문자 인식), 범례 연결, 시각적 검사 및 검증으로 분해하는 다중 에이전트 파이프라인이며, 각 단계별 정보를 추적 가능한 모델 평가 패키지(MEP)로 기록합니다. FinMME 데이터셋에서 AgentFinVQA는 독점 백본을 사용한 기본 모델보다 7.68%p 향상된 성능을 보였습니다 (Gemini-3 Flash; 71.24% vs. 63.56%, McNemar p ≈ 1.1 × 10⁻¹⁶). 또한, 로컬에서 실행되는 오픈 웨이트 모델인 Qwen3.6-27B-FP8을 사용할 때도 4.84%p의 성능 향상을 보였습니다. 검증기의 판단은 유용한 신뢰도 지표로 활용될 수 있으며 (확정된 답변과 수정된 답변에 대한 정확도가 각각 68.2% 및 55.6%), 이를 통해 인간 참여를 통한 리뷰 프로세스를 효율적으로 관리할 수 있습니다. 오류 분석 결과, 질문 이해 부족, 범례 혼동, 추출 오류가 전체 실패의 약 2/3를 차지하며, 이러한 오류는 검증기에 의해 가장 잘 감지되지 않는다는 사실을 알 수 있으며, 이는 향후 연구 방향을 제시합니다. 종합적으로 볼 때, 본 연구 결과는 감사 가능하고 온프레미스 환경에서 실행 가능한 금융 차트 질의응답 시스템이 실용적임을 보여주며, 오픈 웨이트 시스템은 대부분의 정확도 향상을 유지하면서 데이터 보안 및 개인 정보 보호를 보장합니다. 재현 가능한 평가를 위해 관련 코드를 공개합니다.

Original Abstract

Financial chart question answering in regulated settings demands more than accuracy: practitioners must know which answers to trust before acting on them, and many institutions cannot send client data to external model providers. Yet existing chart-QA agents are accuracy-focused and opaque, and most assume proprietary API access; to our knowledge, none combines auditability with on-premise deployability without significant accuracy compromise. We present AgentFinVQA, a multi-agent pipeline that decomposes each query into planning, OCR, legend grounding, visual inspection, and verification, recording every step in a traceable Model Evaluation Packet (MEP) per sample. On FinMME, AgentFinVQA improves $+7.68$ pp over a primary-backbone matched zero-shot baseline with a proprietary backbone (Gemini-3 Flash; 71.24% vs. 63.56%, McNemar $p \approx 1.1 \times 10^{-16}$), and $+4.84$ pp with open-weights Qwen3.6-27B-FP8 served locally. The verifier's verdict also serves as a useful confidence signal (68.2% vs. 55.6% exact accuracy on confirmed vs. revised answers), enabling human-in-the-loop review routing. Error analysis shows that question misunderstanding, legend confusion and extraction error account for nearly two-thirds of failures and are the categories least detected by the verifier, identifying clear directions for future work. Together these results show that auditable, on-premise financial chart QA is practical and that the open-weights system keeps most of the accuracy gains while enabling full data residency. We release our code to support reproducible evaluation.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!