FinAbstain: 불확실성 보정 다중 모드 검색 증강 생성 모델을 활용한 선택적 금융 예측
FinAbstain: Uncertainty-Calibrated Multimodal RAG for Selective Financial Forecasting
대규모 언어 모델(LLM)은 금융 관련 정보를 종합적으로 제시할 수 있지만, 데이터가 부족하거나 오래되었거나 상반되는 경우 높은 확신도를 나타낼 수 있습니다. 이러한 현상은 특히 예측 분야에서 심각한 문제이며, 왜냐하면 보고서, 뉴스, 가격, 거래량 및 기술적 지표 간에 불일치가 발생할 수 있기 때문입니다. 본 연구에서는 불확실성을 고려하여 선택적인 예측을 수행하는 다중 모드 검색 증강 생성(RAG) 프레임워크인 FinAbstain을 제안합니다. FinAbstain은 특정 시점의 정보만을 활용하는 검색 기능을 통해 기본, 뉴스, 기술, 위험 및 검증 담당자에게 해당되는 모달리티에 특화된 정보를 제공합니다. 이들의 확률적 평가는 검색 관련성, 증거의 모순 여부, 반복 샘플 일관성, 그리고 과거 보정 통계와 함께 통합됩니다. 온도 스케일링, 단조 회귀, 컨포멀 예측 및 제안된 하이브리드 불확실성 점수를 포함한 다양한 방법을 동일한 시간 순서 프로토콜에 따라 평가합니다. 제어 시스템은 예측 시 불확실성이 검증된 임계값 이하인 경우에만 강세, 약세 또는 중립적인 결과를 예측하며, 그렇지 않은 경우에는 예측을 보류하거나 추가 증거를 요청하거나 노출을 줄이거나 인간 전문가의 검토를 받도록 합니다. 평가 지표로는 1일 및 5일 비정상 수익 방향, 20일 변동성 구간, 그리고 예측 보류 결정에 대한 정확도, 보정률, 위험-커버리지 비율, 인용 빈도, 거래 성과, 지연 시간 및 비용 등을 사용합니다. 실제 데이터 수집이 완료되기 전에 설계의 투명성을 확보하기 위해, 우리는 실제 실험 결과 대신 명시적으로 라벨링된 시뮬레이션 결과를 보고합니다. 이러한 결과는 다음과 같은 가설을 뒷받침합니다: 보정된 예측 보류는 높은 정확도를 얻기 위해 커버리지를 희생할 수 있습니다. 본 연구의 기여점은 시간 안전 아키텍처, 통합된 불확실성 공식 및 증거 기반 선택적 금융 예측을 위한 재현 가능한 평가 방법론입니다.
Large language models (LLMs) can synthesize financial narratives but may express high confidence when evidence is sparse, stale, or contradictory. This failure is especially consequential in forecasting, where filings, news, prices, volume, and technical signals can disagree. We present FinAbstain, a research framework for uncertainty-calibrated multimodal retrieval-augmented generation (RAG) with selective prediction. A point-in-time retriever admits only information public at the forecast timestamp and supplies modality-specific evidence to fundamental, news, technical, risk, and verification agents. Their probabilistic assessments are aggregated with retrieval relevance, evidence contradiction, repeated-sample consistency, and historical calibration statistics. Temperature scaling, isotonic regression, conformal prediction, and a proposed hybrid uncertainty score are evaluated under a common chronological protocol. A controller predicts bullish, bearish, or neutral outcomes only when uncertainty is below a validated threshold; otherwise it abstains, requests evidence, reduces exposure, or routes the case to human review. The evaluation covers one- and five-day abnormal-return direction, twenty-day volatility intervals, and abstention decisions, using accuracy, calibration, risk--coverage, citation, trading, latency, and cost metrics. To make the design auditable before a full data collection is complete, we report explicitly labeled simulated results rather than empirical claims. These results illustrate the intended hypothesis: calibrated abstention may trade coverage for lower selective error and drawdown. The contribution is a time-safe architecture, a composite uncertainty formulation, and a reproducible evaluation blueprint for evidence-grounded selective financial forecasting.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.