Omni-Decision: 다양한 형태의 데이터를 활용한 질문 답변 시스템을 위한 증거 기반 에이전트
Omni-Decision: A Progressive Evidence-State Agent System for Omni-Modal QA
다중 모달(Multi-Modal) 질의 응답(QA)은 비디오, 오디오, 이미지, 웹 페이지 및 계산 결과 등 여러 곳에 흩어져 있는 정보를 바탕으로 질문에 답해야 하는 시스템을 필요로 합니다. 기존의 다중 모달 에이전트 시스템들은 종종 중간 과정에서 생성된 정보들을 임시 저장 공간, 도구 사용 기록 또는 자유 형식의 이력 데이터로 남겨두기 때문에, 어떤 정보가 확인되었고, 어떤 정보가 누락되었으며, 언제 답변에 필요한 충분한 증거를 확보했는지 추적하기 어렵습니다. 본 논문에서는 학습 과정이 필요 없는 증거 상태 시스템인 Omni-Decision을 제안합니다. Omni-Decision은 다중 모달 질의 응답을 특정 질문 범위 내에서 해결 가능한 증거 수집 프로세스로 변환합니다. 각 질문에 대해, Omni-Decision은 확인된 증거, 해결되지 않은 충돌, 사실 및 계산 의존성, 그리고 필요한 증거 목록을 포함하는 구조화된 증거 상태를 유지합니다. 공유된 상태 정보는 계획 수립, 증거 획득, 검증, 수정 및 최종 단계에 영향을 미칩니다. 미디어, 웹, 계산 및 검증 모듈에서 얻은 다양한 정보를 표준화하고, 판단하여 결정적인 상태 업데이트를 통해 시스템에 반영합니다. 이러한 설계는 목표 지향적인 증거 획득을 가능하게 하고, 희소한 데이터 간의 연관성을 유지하며, 수정 과정과 종료 조건에 대한 투명한 제어를 제공합니다. Omni-Decision은 OmniGAIA 데이터셋에서 45.6%의 정확도를, WorldSense 데이터셋에서 58.3%의 정확도를 달성하여 기존 모델 대비 각각 +27.3%, +30.2%의 성능 향상을 보였습니다. 증거 상태 제어 기능을 제거한 실험 및 실행 경로 분석 결과는 다단계 다중 모달 정보 수집 과정에서 명시적인 증거 상태 제어가 중요한 역할을 한다는 것을 뒷받침합니다.
Omni-modal evidence-seeking QA requires agents to answer questions whose evidence is sparsely distributed across videos, audio, images, web pages, and computation results. Existing agentic multimodal systems often leave evidence in scratchpads, tool trajectories, or free-form histories, making it difficult to track what has been grounded, what remains missing, and when the evidence is sufficient to answer. We propose Omni-Decision, a training-free evidence-state system that turns omni-modal QA into a query-scoped evidence-closure process. For each query, Omni-Decision maintains a structured evidence state containing confirmed evidence, unresolved conflicts, fact and computation dependencies, and open evidence needs. A shared state view conditions planning, evidence acquisition, validation, repair, and finalization. Heterogeneous observations from media, web, computation, and verification modules are normalized, judged, and committed through deterministic state updates. This design enables targeted evidence acquisition, preserves sparse cross-modal cues, and provides inspectable control over repair and stopping. Omni-Decision achieves 45.6% accuracy on OmniGAIA and 58.3% on WorldSense, improving over the baselines by +27.3 and +30.2 percentage points, respectively. No-state ablations and trajectory audits further support the role of explicit evidence-state control in multi-step omni-modal evidence seeking.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.