2605.28683v1 May 27, 2026 cs.AI

VeriTrip: 비정형 웹 데이터 기반 여행 계획 에이전트에 대한 검증 가능한 성능 평가 도구

VeriTrip: A Verifiable Benchmark for Travel Planning Agents over Unstructured Web Corpora

Jian Liang
Jian Liang
Citations: 9,396
h-index: 5
Mu Xu
Mu Xu
Citations: 82
h-index: 6
Hang Zhang
Hang Zhang
Citations: 198
h-index: 5
Jiayi Tian
Jiayi Tian
Citations: 195
h-index: 5
Xin Xiong
Xin Xiong
Citations: 1
h-index: 1
Xiaoyun Zhang
Xiaoyun Zhang
Citations: 2,827
h-index: 2
Yuting Xu
Yuting Xu
Citations: 431
h-index: 8

기존의 성능 평가 도구들은 API 중심적인 방식으로 여행 계획 에이전트를 평가하는 데 중요한 역할을 해왔습니다. 그러나 자율 에이전트의 기능이 발전함에 따라, 단순한 도구 실행을 넘어 개방형 웹 환경의 복잡성을 처리할 수 있도록 평가 방법은 진화해야 합니다. 현재의 성능 평가 도구들은 정보 노이즈를 고려하지 않거나, 다양한 출처 간 사실상의 모순을 무시하고, 시각적 인지 능력을 논리적인 계획에 통합하는 필요성을 간과합니다. 본 연구에서는 에이전트의 견고성과 신뢰성에 대한 증가하는 요구 사항을 충족하기 위해 설계된 검증 가능한 성능 평가 도구인 VeriTrip을 소개합니다. VeriTrip은 비정형 멀티모달 웹 데이터 기반의 증거에 근거한 추론 능력을 평가하는 데 중점을 둡니다. 실제 출처에서 파생된 멀티모달 검색 기반(MRB)을 구축하여, 에이전트가 다양한 유형의 데이터를 자율적으로 처리하도록 합니다. 동기화된 검증 가능한 지식 기반(VKB)은 셀 단위의 검증 프로토콜을 통해 사실적인 신뢰도를 정확하게 측정하고, 체계적인 추론 실패와 매개변수 환각 현상을 구별합니다. 선도적인 멀티모달 대규모 언어 모델(MLLM)에 대한 평가 결과, 자율적인 검색 과정이 인지적 부담을 증가시켜 명령어 기억력을 저하시키는 중요한 '검색-추론 균형' 문제가 있음을 확인했습니다. VeriTrip은 제약 없이 다양한 유형의 환경에서 작동할 수 있는 차세대 계획 에이전트를 개발하는 데 필요한 엄격한 기반을 제공합니다.

Original Abstract

Existing benchmarks have laid the foundation for travel planning agents by establishing API-centric paradigms. However, as the capabilities of Autonomous Agents continue to advance, their evaluation must evolve beyond simple tool execution toward handling the inherent complexities of the open web. Current benchmarks bypass core cognitive hurdles: they fail to account for information noise, ignore multi-source factual contradictions, and overlook the necessity of grounding visual perception into logical planning. We introduce VeriTrip, a verifiable benchmark designed to meet the increasing demands for agent robustness and reliability. VeriTrip shifts the evaluation focus to evidence-grounded reasoning over unstructured multimodal web corpora. It establishes a Multimodal Retrieval Base (MRB) derived from real-world sources, forcing agents to autonomously orchestrate queries across heterogeneous data. A synchronized Verifiable Knowledge Base (VKB) enables a cell-wise verification protocol that precisely quantifies factual reliability, distinguishing systematic reasoning failures from parametric hallucinations. Our evaluations across leading MLLMs reveal a critical \textit{retrieval-reasoning trade-off}: the cognitive load of autonomous retrieval significantly erodes instruction retention. VeriTrip provides the rigorous foundation necessary for the next generation of planning agents capable of operating in unconstrained, multimodal environments.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!