2605.27164v1 May 26, 2026 cs.AI

질의를 기호적으로 수행할 것인가, 아니면 의미론적으로 검색할 것인가? 준정형 질문 응답을 위한 데이터셋 및 방법

Query Symbolically or Retrieve Semantically? A Dataset and Method for Semi-Structured Question Answering

Mateusz Czyznikiewicz
Mateusz Czyznikiewicz
Citations: 3
h-index: 1
Ryszard Tuora
Ryszard Tuora
Citations: 20
h-index: 2
A. Kozakiewicz
A. Kozakiewicz
Citations: 4
h-index: 1
Mateusz Gali'nski
Mateusz Gali'nski
Citations: 0
h-index: 0
M. Godziszewski
M. Godziszewski
Citations: 55
h-index: 4
Michael A. Karpowicz
Michael A. Karpowicz
Citations: 281
h-index: 5
Cristina Cornelio
Cristina Cornelio
Citations: 32
h-index: 3
T. Hospedales
T. Hospedales
Citations: 3
h-index: 1
Tomasz Zietkiewicz
Tomasz Zietkiewicz
Citations: 0
h-index: 0

질문 응답을 위한 검색 증강 생성(Retrieval-Augmented Generation, RAG) 시스템은 일반적으로 질의와 문서 조각 간의 의미 유사성을 기반으로 증거를 검색합니다. 이 접근 방식은 비정형 텍스트에는 효과적이지만, 답변에 정확한 필터링, 집계 또는 여러 문서에 걸쳐 구조화된 속성에 대한 완전한 검색이 필요한 준정형 데이터 코퍼스에서는 신뢰성이 떨어집니다. 기호 기반 접근 방식은 이러한 작업을 지원하지만, 노이즈가 많은 자연어 코퍼스에서는 종종 불안정합니다. 우리는 텍스트 지식 그래프를 사용하여 의미론적 검색을 수행하고, 기호 지식 그래프를 사용하여 유형화된 주체-술어-객체 삼중항에 대한 기호 질의를 수행하는 RAG 프레임워크인 DualGraph로 이러한 격차를 해소합니다. 이 두 가지 구성 요소를 기반으로, 의미론적 및 기호 증거를 선택하거나 결합하기 위한 다양한 전략을 제공합니다. 또한, 우리는 준정형 제품 문서와 광범위한 개방형 및 사양 지향 검색 질문이 수동으로 큐레이팅된 상업용 쇼핑 웹사이트에서 가져온 벤치마크인 SpecsQA를 소개합니다. 실험 결과, DualGraph는 다양한 유형의 질문에 대해 최첨단 밀집 검색(dense-retrieval), GraphRAG, 기호 기반 및 테이블 지향 기준 모델보다 일관되게 우수한 성능을 보였습니다. 코드와 데이터는 https://github.com/corneliocristina/DualGraphRAG에서 확인할 수 있습니다.

Original Abstract

Retrieval-Augmented Generation (RAG) systems for question answering typically retrieve evidence by semantic similarity between the query and document chunks. While effective for unstructured text, this approach is less reliable on semi-structured corpora where answering may require exact filtering, aggregation, or exhaustive retrieval over structured attributes across multiple documents. Symbolic approaches support such operations, but they are often brittle on noisy natural-language corpora. We address this gap with DualGraph, a RAG framework that represents documents through two complementary views: a Textual Knowledge Graph for semantic retrieval and a Symbolic Knowledge Graph for symbolic querying over typed subject--predicate--object triples. Building on these two components, we provide multiple strategies for selecting or combining semantic and symbolic evidence.We also introduce SpecsQA, a benchmark from a commercial shopping website with semi-structured product documents and manually curated questions spanning open-ended and specification-oriented retrieval. Experiments show that DualGraph consistently outperforms state-of-the-art dense-retrieval, GraphRAG, symbolic, and table-oriented baselines across question types.Code and data are available at https://github.com/corneliocristina/DualGraphRAG.

0 Citations
0 Influential
25.9657359028 Altmetric
0.0 Score
Original PDF
1

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!