2605.29250v1 May 28, 2026 cs.CL

OmniRetrieval: 이질적인 지식 소스 간의 통합 검색

OmniRetrieval: Unified Retrieval across Heterogeneous Knowledge Sources

Patara Trirat
Patara Trirat
KAIST
Citations: 261
h-index: 6
Heejun Lee
Heejun Lee
Citations: 69
h-index: 6
S. Hwang
S. Hwang
Citations: 92
h-index: 3
Sangwoo Park
Sangwoo Park
Citations: 13
h-index: 3
Minki Kang
Minki Kang
Citations: 923
h-index: 15
Woongyeong Yeo
Woongyeong Yeo
KAIST
Citations: 407
h-index: 5
Jinheon Baek
Jinheon Baek
Citations: 2,712
h-index: 21
Soyeong Jeong
Soyeong Jeong
Citations: 1,002
h-index: 11

실제 정보 요구 사항은 비정형 텍스트, 관계형 테이블부터 지식 그래프 및 속성 그래프에 이르기까지 구조적으로 다양한 지식 소스에 대한 접근을 필요로 합니다. 기존 검색 시스템들은 일반적으로 고정된 질의 언어 하에서 한 번에 하나의 소스만 처리하며, 이는 다양한 지식 소스를 호환되지 않는 인터페이스로 인해 분산시키는 문제를 야기합니다. 자연스러운 통합 시도는 이러한 소스들을 공유 공간으로 묶는 것이지만, 이는 각 소스가 가진 표현력을 제공하는 구조적 특징(예: 스키마, 온톨로지, 합성 연산자)을 무시하게 됩니다. 따라서 다양한 지식에 대한 효과적인 검색은 동질화가 아닌, 각 소스의 특성을 존중하면서 이를 통합할 수 있는 상위 계층이 필요합니다. 이러한 목표를 달성하기 위해, 우리는 OmniRetrieval이라는 프레임워크를 제안합니다. OmniRetrieval은 모든 자연어 질의를 받아들이고, 적절한 지식 소스를 식별하며, 해당 소스에 맞는 질의를 해당 실행 엔진으로 전송합니다. 텍스트, 관계형 및 그래프 구조 데이터로 구성된 13개의 데이터 세트와 309개의 다양한 지식 베이스를 포함하는 광범위한 벤치마크에서 OmniRetrieval은 단일 소스 기반 시스템보다 우수한 성능을 보였으며, 이는 OmniRetrieval이 이질적인 소스를 위한 범용 인터페이스 역할을 하면서 각 소스의 가치를 유지하는 데 효과적임을 보여줍니다.

Original Abstract

Real-world information needs require access to structurally diverse knowledge sources, from unstructured text and relational tables to knowledge graphs and property graphs. Existing retrievers, however, operate over one source at a time under a fixed query language, leaving the broader landscape of available knowledge fragmented behind incompatible interfaces. A natural attempt at unification would collapse these sources into a shared space, but this erases the structural affordances (such as schemas, ontologies, compositional operators) that give each source its expressive power. Effective retrieval over diverse knowledge, therefore, requires not homogenization but an overarching layer that meets each source on its own terms. To achieve this, we present OmniRetrieval, a framework that takes any natural-language query, identifies appropriate knowledge sources, and dispatches source-native queries to their native execution engines. Across an extensive benchmark spanning 13 datasets and 309 distinct knowledge bases over text, relational, and graph-structured sources, OmniRetrieval exceeds single-source baselines, demonstrating that it can serve as a general-purpose interface to the heterogeneous sources while preserving the structural distinctions that make each source valuable.

0 Citations
0 Influential
10.5 Altmetric
52.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!