2607.24063v1 Jul 27, 2026 cs.AI

알고 얻는 비용: 정적 순위표를 넘어선 환각 현상 평가를 위한 자원 기반 프로토콜

The Cost of Knowing: A Resource-Aware Protocol for Benchmarking Hallucination Beyond Static Leaderboards

Keyu Li
Keyu Li
Citations: 42
h-index: 3
Dequan Wang
Dequan Wang
Citations: 17
h-index: 1
Jin Gao
Jin Gao
Citations: 223
h-index: 5

현재 최첨단 모델들은 사실성 평가에서 최고 수준에 근접하게 분포하고 있습니다. 따라서 중요한 질문은 시스템의 사실성이 얼마나 많은 컴퓨팅 자원을 소비하는가로 옮겨가고 있습니다. 기존의 정적 순위표는 사실성을 개별적으로 평가하며, 컴퓨팅 자원을 무시하기 때문에 진정으로 더 나은 시스템과 단순히 더 많은 자원을 사용하는 시스템을 구분할 수 없습니다. 예를 들어, '최고 4개' 모델 중 가장 성능이 좋은 모델이 원본 점수(H-Score)는 높지만(0.9169 대 0.9103), 비용을 고려하면 더 낮은 Q-Score(0.5169 대 0.5217)를 기록하며, 약 네 배의 토큰 사용량과 지연 시간을 갖습니다. 따라서 정적 순위표에서 가장 높은 점수를 받는 시스템이 실제로는 배포하기에 더 나쁜 시스템일 수 있습니다. 이러한 상충 관계를 명확하게 하기 위해, 우리는 자원 기반 평가 프로토콜인 MAS-HQ(Multi-Agent System Hallucination Quest)를 소개합니다. MAS-HQ는 모든 사실성 검사기를 감싸고 비용을 정규화하며, 시스템들을 개별적으로 평가하는 대신 서로 경쟁하도록 합니다. Q-Score는 경쟁적인 환경에서 측정된 사실성에서 정규화된 비용을 뺀 값을 나타냅니다. 요약 및 개방형 질의응답 분야에서 단일 에이전트 기반 모델은 자원 집약적인 과최적화로 이어지는 경향이 있는 반면, 경쟁 환경에서는 더 효율적인 정책이 도출됩니다. 이러한 이점은 작지만 일관되며, 100번의 실험을 통해 안정적으로 나타났습니다. MAS-HQ는 최첨단 시스템(Gemini-2.5-Pro 및 시뮬레이션된 GPT-5)의 원본 사실성 점수가 이미 최고점에 근접한 경우에도 차별성을 유지합니다. MAS-HQ는 사실적인 답변을 얻는 데 드는 비용을 측정하는 재현 가능한 방법을 제공합니다.

Original Abstract

On standard factuality tasks, frontier models now cluster near the top of the scale. The question is therefore shifting from how factual a system is toward how much compute that factuality costs. Static leaderboards score factuality in isolation and treat compute as free, so they cannot tell a genuinely better system apart from one that simply spends more. Consider a ranking reversal. A brute-force Best-of-4 agent posts the higher raw factuality score (H-Score 0.9169 vs 0.9103) and would top a static leaderboard, but once cost is counted it is the worse system, losing on Q-Score (0.5169 vs 0.5217) at roughly four times the tokens and latency, under a reported cost weight whose sensitivity we sweep. So the system that tops a static leaderboard can be the worse one to deploy. To make this trade-off visible, we introduce MAS-HQ (Multi-Agent System Hallucination Quest), a resource-aware evaluation protocol. It wraps any factuality detector and normalizes for cost, and it pits systems against each other rather than scoring them in isolation. The Q-Score measures factuality minus normalized cost under a competitive match. Across summarization and open-domain QA, single-agent baselines drift into resource-heavy over-optimization, while competition elicits more resource-efficient policies. These gains are small but consistent, and stable across 100 trials. The axis stays discriminative for frontier systems (Gemini-2.5-Pro, and GPT-5) whose raw factuality scores are already bunched near the ceiling. MAS-HQ provides a reproducible way to measure how much a factual answer costs.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!