2607.17742v1 Jul 20, 2026 cs.AI

의미적으로 유사하지만 논리적으로 구별되는 테이블 RAG에서의 의미-답변 가능성 격차 진단

Semantically Similar, Logically Distinct: Diagnosing the Semantic-Answerability Gap in Table RAG

Haobo Wang
Haobo Wang
Citations: 1,249
h-index: 17
Zujie Ren
Zujie Ren
Citations: 3
h-index: 1
Wen-song Ye
Wen-song Ye
Citations: 314
h-index: 7
Gang Chen
Gang Chen
Citations: 79
h-index: 2
Jiaming Tian
Jiaming Tian
Citations: 11
h-index: 2
Liyao Li
Liyao Li
Citations: 176
h-index: 3
Lihua Yu
Lihua Yu
Citations: 59
h-index: 4
Junbo Zhao
Junbo Zhao
Citations: 97
h-index: 4

테이블은 검색 증강 생성(RAG)에서 중요한 지식 출처이지만, 검색된 테이블이 쿼리에 대한 답변을 제공하기에 충분한 정보를 포함하지 못하는 경우가 있습니다. 이를 우리는 '답변 가능성'이라고 부릅니다. 답변 가능성은 일반적으로 소스 또는 여러 소스가 충분한 증거를 포함하고 있는지 여부에 관한 광범위한 개념입니다. 그러나 의미적 관련성을 최적화하도록 설계된 검색 모델은 단일 소스의 경우에도 이를 보장하지 못하며, 이는 근본적인 불일치를 야기합니다. 이러한 문제를 연구하기 위해, 우리는 테이블 콘텐츠 수준의 답변 가능성에 대한 진단 벤치마크인 TCR-Bench를 개발했습니다. TCR-Bench는 매우 유사한 스키마를 갖지만 미묘한 내용 차이가 있는 '형제 테이블'을 기반으로 합니다. TCR-Bench에서 평가된 밀집 검색 모델은 지속적으로 '의미-답변 가능성 격차'를 보입니다. 즉, 올바른 형제 그룹을 검색하는 데는 성공하지만, 그 안에서 답변 가능한 특정 테이블을 정확히 식별하는 데 어려움을 겪으며, 질문 응답 성능이 0.755 (oracle)에서 0.330 (상위 5개 검색 결과)으로 떨어집니다. 분석 결과, 이 격차는 의미적 누적, 스키마 수준의 단서 의존성 및 약한 행-열 결합과 관련되어 있음을 알 수 있습니다. 이러한 격차의 원인을 진단하기 위해, 우리는 경량화된 2단계 파이프라인인 Answerability-Aware Reranking (AAR)을 사용하여 직접적인 쿼리-테이블 답변 가능성 판단을 적용하여 성능을 향상시킬 수 있는지 테스트했습니다. AAR은 상위 1개 대상 테이블 검색률을 18.2%에서 57.4%로 크게 향상시켰습니다. 이러한 큰 개선은 관찰된 실패의 상당 부분이 모델 용량의 고유한 제한이 아닌, 누락된 답변 가능성 검증 단계에 기인한다는 증거입니다.

Original Abstract

Tables are a critical knowledge source in retrieval-augmented generation (RAG), but a retrieved table may lack sufficient evidence to answer a query, a property we call answerability. While answerability broadly concerns whether a source or collection of sources contains sufficient evidence, retrieval models optimized for semantic relevance do not guarantee it even in the single-source case, creating a fundamental mismatch. To study this, we introduce TCR-Bench, a diagnostic benchmark for Table Content-level Answerability in RAG, built around sibling tables, i.e., tables with highly similar schemas but subtle content differences. On TCR-Bench, the dense retrievers we evaluate persistently exhibit a Semantic-Answerability Gap: they often retrieve the correct sibling group yet struggle to pinpoint the uniquely answerable table within it, dropping QA performance from 0.755 (oracle) to 0.330 (top-5 retrieved). Our analysis suggests this gap is associated with semantic accumulation, schema-level cue dependence, and weak row-column binding. As a diagnostic probe into the source of this gap, we test whether a lightweight two-stage pipeline, Answerability-Aware Reranking (AAR), applying direct query-table answerability judgment, can recover performance: it raises top-1 target retrieval from 18.2% to 57.4%, and this large gain is itself evidence that much of the observed failure reflects a missing answerability verification step, rather than an inherent limitation of model capacity alone.

0 Citations
0 Influential
8.5 Altmetric
42.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!