2606.31800v1 Jun 30, 2026 cs.AI

Evo-PI: 원칙 기반의 능동적 지도를 통한 의료 추론 정렬

Evo-PI: Aligning Medical Reasoning via Evolving Principle-Guided Supervision

M. Witbrock
M. Witbrock
Citations: 11
h-index: 2
Xianda Zheng
Xianda Zheng
Citations: 162
h-index: 5
Huan Gao
Huan Gao
Citations: 261
h-index: 7
Meng-Fen Chiang
Meng-Fen Chiang
Citations: 99
h-index: 5
Kaiqi Zhao
Kaiqi Zhao
Citations: 1
h-index: 1
Shangyang Li
Shangyang Li
Citations: 96
h-index: 6

최근 상당한 발전에도 불구하고, 대규모 다중 모드 언어 모델(MLLM)의 추론 능력은 여전히 고정된 지도 방식에 의해 근본적으로 제약됩니다. 이러한 고정된 프롬프트, 규칙 또는 보상 모델은 학습 과정 전반에 걸쳐 비적응적인 지침을 제공하며, 이는 출력 형식을 강제하는 데는 충분하지만, 근본적인 추론 과정을 형성하지 못하여 복잡한 의사 결정 작업에서 일반화 능력 저하 및 성능 포화를 초래합니다. 본 논문에서는 추론 원칙을 명시적인 언어 기반의 지도 신호로 간주하고 생성, 평가 및 반복적으로 발전시키는 원칙 중심 학습 프레임워크인 Evo-PI를 제안합니다. Evo-PI는 고정된 보상에 의존하는 대신, 모델의 추론을 안내하는 원칙과 모델의 행동이 이러한 원칙을 개선하는 상호 진화 루프를 가능하게 합니다. 이 동적인 정렬 메커니즘은 지도가 모델의 추론 결점을 점진적으로 적응하도록 만듭니다. 우리는 Evo-PI를 의료 시각 질의 응답에 적용하여 구조화된 시각-텍스트 추론을 요구하는 고위험 테스트 환경으로 사용했습니다. 8개의 벤치마크와 다양한 모델 아키텍처에서 Evo-PI는 일관되게 추론 정확도를 향상시키며, 최대 24.6%의 성능 향상을 달성했습니다. 이러한 결과는 원칙 기반의 능동적 지도가 MLLM에서 전문가 수준의 추론을 학습시키는 확장 가능하고 일반적인 패러다임을 제공한다는 것을 시사합니다. 코드: https://github.com/zhengxianda/Evo_PI.

Original Abstract

Despite recent progress, the reasoning capabilities of large multimodal language models (MLLMs) remain fundamentally constrained by static supervision, where fixed prompts, rules, or reward models provide non-adaptive guidance throughout training. Such static signals are often sufficient to enforce output formats, but fail to shape the underlying reasoning process, leading to brittle generalization and performance saturation in complex decision-making tasks. We propose Evo-PI, a principle-centric learning framework that treats reasoning principles as explicit, language-based supervision signals that can be generated, evaluated, and iteratively evolved. Instead of relying on fixed rewards, Evo-PI enables a co-evolutionary loop in which principles guide model reasoning, while model behaviors in turn refine the principles that supervise them. This dynamic alignment mechanism allows supervision to progressively adapt to the model's reasoning deficiencies. We instantiate Evo-PI in medical visual question answering as a high-stakes testbed requiring structured visual-textual reasoning. Across eight benchmarks and multiple model backbones, Evo-PI consistently improves reasoning accuracy, achieving gains of up to 24.6%. Our results suggest that evolving principle-guided supervision offers a scalable and general paradigm for training expert-aligned reasoning in MLLMs. Code is available at https://github.com/zhengxianda/Evo_PI.

0 Citations
0 Influential
26.9657359028 Altmetric
0.0 Score
Original PDF
1

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!