2605.26691v1 May 26, 2026 cs.AI

도구 오류에 유의하십시오: 의료 분야 AI 에이전트의 시너지 효과 극대화를 위한 도구 활용

Mind the Tool Failures: Achieving Synergistic Tool Gains for Medical Agents

Kaiyu Guo
Kaiyu Guo
Citations: 13
h-index: 2
Tan Pan
Tan Pan
Citations: 34
h-index: 4
Limei Han
Limei Han
Citations: 33
h-index: 3
Yuan Cheng
Yuan Cheng
Citations: 31
h-index: 3
Yu Gan
Yu Gan
Citations: 15
h-index: 2
Weimiao Yu
Weimiao Yu
Citations: 49
h-index: 3
Guangnan Ye
Guangnan Ye
Citations: 1
h-index: 1
Chenzhi Jiang
Chenzhi Jiang
Citations: 0
h-index: 0

최근 의료 AI 에이전트는 진단, 치료 추천 및 증거 검색을 위해 외부 도구를 점점 더 많이 사용하고 있지만, 대부분의 기존 방법은 작업에 적합한 도구가 의도된 범위 내에서 신뢰할 수 있다고 가정합니다. 그러나 실제 임상 환경에서는 관련성이 높은 도구조차도 어려운 경우에 실패하여 안전하지 않은 결과를 초래할 수 있습니다. 이러한 문제를 해결하기 위해, 우리는 개별 도구가 놓치는 오류 사례를 수정하기 위해 불완전한 도구 환경에서의 의료 도구 사용을 연구합니다. 인스턴스 의존적인 오류 패턴은 최적의 단일 도구와 이상적인 인스턴스 기반 선택기 사이의 격차를 만듭니다. 이 격차는 'Single-Oracle 위험 간극'이라고 불립니다. 핵심 과제는 기존의 작업 수준 도구 선택이 이러한 격차를 실현할 수 없다는 점인데, 이는 최상의 단일 도구의 성능에 의해 본질적으로 제한되기 때문입니다. 이러한 관찰에 따라, 우리는 인스턴스 수준의 이질성을 고려하고 도구 사용을 인스턴스 수준의 선택 문제로 공식화합니다. 특히, 우리는 확률적 위험 최소화 및 불일치 인식 시너지 학습을 위한 보상을 제공하는 GRPO 기반 강화 학습 프레임워크를 제안하며, 이를 통해 오류가 있는 도구 합의를 인스턴스 수준에서 수정할 수 있습니다. 또한, 엔트로피 지향 샘플링 전략을 채택하여 학습에 더 강력한 신호를 제공하는 불일치가 높은 인스턴스의 중요도를 높입니다. 이 두 가지 요소는 서로 보완되어 인스턴스 수준의 이질성을 완화하고 도구 시너지 효과를 향상시킵니다. 두 가지 작업 및 일곱 개의 의료 벤치마크에 대한 실험 결과, 제안된 방법은 다양한 기본 모델보다 지속적이고 안정적인 성능 향상을 보여주며, 신뢰할 수 있는 의료 AI 시스템을 위한 시너지 기반 도구 활용의 중요성을 강조합니다.

Original Abstract

Medical AI agents increasingly use external tools for diagnosis, treatment recommendation, and evidence retrieval, yet most existing approaches assume that task-appropriate tools are reliable within their intended scope. This assumption is fragile in real clinical settings, where even relevant tools may fail on challenging instances and lead to unsafe downstream decisions. To address this issue, we study medical tool use under imperfect-tool settings to correct failure instances missed by individual tools. Instance-dependent failure patterns create a gap between the best fixed single tool and an ideal instance-wise selector, which we refer to as the Single-Oracle risk gap. The core challenge is that conventional task-level tool selection cannot realize this gap, as it is inherently bounded by the performance of the best single tool. Motivated by this observation, we therefore account for instance-level heterogeneity and formulate tool use as an instance-level selection problem. Particularly, we propose a GRPO-based reinforcement learning framework with rewards for probabilistic risk minimization and disagreement-aware synergy learning, which promotes instance-level correction of erroneous tool consensus. Furthermore, an entropy-guided sampling strategy is adopted to upweight high-disagreement instances, which provide stronger signals for learning instance-specific tool synergy. These two components complement each other in mitigating instance-level heterogeneity and improving tool synergy. Experiments on two tasks and seven medical benchmarks show that our method consistently achieves robust and stable improvements over a broad range of baselines, highlighting the importance of synergy-aware tool use for reliable medical agentic systems.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!