2607.28225v1 Jul 30, 2026 cs.CV

FaithEyes: 멀티 에이전트 프로세스-이미지 검증을 통한 신뢰성 있는 도구 사용 연구

FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification

Yehui Tang
Yehui Tang
Citations: 18
h-index: 3
Wei Xia
Wei Xia
Citations: 82
h-index: 5
Ziheng Li
Ziheng Li
Citations: 20
h-index: 2
Xingrun Xing
Xingrun Xing
Citations: 594
h-index: 8
Haoqing Wang
Haoqing Wang
Citations: 582
h-index: 8

텍스트 추론과 이미지 자르기, 코드 기반 이미지 조작 등 명시적인 도구 호출을 결합하여 신뢰성과 해석 가능성을 높이는 시각-언어 모델(VLMs)이 주목받고 있습니다. 하지만 최근 연구에 따르면 이러한 모델들이 종종 도구를 잘못 사용하는 경우가 있다는 것이 밝혀졌습니다. 많은 경우, 생성된 이미지가 질문과 관련이 없는데도 (예: 도구가 잘못된 영역을 자르거나 검색 대상 객체를 놓치는 경우), 모델은 여전히 정답을 제공하며, 해당 도구 호출에 대한 보상을 받습니다. 이러한 불필요하거나 부정확한 도구 호출은 계산 자원을 낭비하며, 모델이 실제로 수집한 증거보다는 사전 지식이나 원본 이미지에 의존하고 있음을 드러냅니다. 이는 기존 방법의 두 가지 한계에서 비롯됩니다. 첫째, 도구 보상이 유용한 호출과 무용지물인 호출을 구별하지 못한다는 점이며, 둘째, 도구 피드백이 유용성에 대한 정보를 담고 있지 않다는 점입니다. 이러한 문제를 해결하기 위해, 저희는 멀티 에이전트 자기 평가 프레임워크인 FaithEyes를 제안합니다. 구체적으로, VLM을 사용하여 각 프로세스 이미지가 질문에 답변하는 데 도움이 되는지 판단합니다. 이 판단은 추론 과정에서 도구 관찰의 일부로 주입되어 후속 추론을 돕고, 동시에 유용한 도구 비율에 따라 도구 보상을 조정하여 잘못된 보상 습득을 방지합니다. 또한, 평가 시에도 판단 정보를 활용하여 학습과 테스트 간의 일관성을 유지하기 위해, 모델 자체를 하위 에이전트로 사용하여 주 에이전트가 수행한 도구 호출을 평가하는 멀티 에이전트 프레임워크를 설계했습니다. FaithEyes는 수정된 공개 데이터셋에 대한 두 단계의 SFT + RL 파이프라인을 통해 학습되었으며, 시각적 인식 및 추론 벤치마크에서 경쟁력 있는 또는 우수한 정확도를 달성하는 동시에 도구 사용의 신뢰성을 크게 향상시켰습니다. 자세한 내용은 다음 홈페이지에서 확인할 수 있습니다: https://github.com/Mosi-AI/FaithEyes.

Original Abstract

Agentic vision-language models (VLMs), which interleave textual reasoning with explicit tool calls such as cropping and code-based image manipulation, have emerged as a compelling paradigm for reliable and interpretable multimodal reasoning. However, recent studies have revealed that such models often use tools unfaithfully. Many process images are irrelevant to the question (e.g., the tool crops the wrong region or misses the queried target), yet the call still receives full credit and the model still answers correctly. Such decorative or misaligned tool calls waste computation and reveal that the model leans on prior knowledge or the original image rather than the evidence it retrieves. This may stem from two limitations of prevailing methods: the tool reward fails to distinguish useful from useless calls, and tool feedback carries no signal of usefulness. To this end, we introduce FaithEyes, a multi-agent self-judging framework. Concretely, we use a VLM to judge whether each process image helps answer the question. The judgement is injected into the reasoning context as part of the tool observation to help subsequent reasoning, and meanwhile is used to scale the tool reward by the helpful-tool ratio to suppress reward hacking. To keep judgement available at evaluation and thus ensure train-test consistency, we further design a multi-agent framework where the model itself serves as a subagent to judge the tool calls from main agent, eliminating any dependence on an external model at inference. Training via a two-stage SFT + RL pipeline on adapted open-source data, FaithEyes attains competitive or superior accuracy across visual perception and reasoning benchmarks, while markedly improving tool faithfulness. The homepage is at https://github.com/Mosi-AI/FaithEyes.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!