2606.09052v1 Jun 08, 2026 cs.LG

INFUSER: 영향력 기반 자기 진화가 추론 능력을 향상시킨다

INFUSER: Influence-Guided Self-Evolution Improves Reasoning

Shuang Li
Shuang Li
Citations: 84
h-index: 2
Fengzhuo Zhang
Fengzhuo Zhang
Citations: 233
h-index: 5
Zhuoran Yang
Zhuoran Yang
Citations: 56
h-index: 3
Siyu Chen
Siyu Chen
Citations: 194
h-index: 6
Miao Lu
Miao Lu
Citations: 20
h-index: 3
Beining Wu
Beining Wu
Citations: 38
h-index: 4
Heejune Sheen
Heejune Sheen
Citations: 151
h-index: 5
Zhiyuan Li
Zhiyuan Li
Citations: 14
h-index: 2
Jose Blanchet
Jose Blanchet
Citations: 20
h-index: 1
Tianhao Wang
Tianhao Wang
Citations: 8
h-index: 2

자기 진화는 강력한 추론 능력 확보를 위한 확장 가능한 방법입니다. 사전 학습된 언어 모델은 최소한의 외부 감독만으로 스스로를 개선할 수 있습니다. 그러나 기존 방법들은 광범위하게 선별된 또는 교사(teacher)가 생성한 훈련 데이터에 의존하거나, 생성기가 비지도 방식으로 실행될 때, 솔버(solver)를 실제로 향상시키지 못하는 난이도 기반 보상을 제공합니다. 우리는 INFUSER라는 반복적인 공동 훈련 프레임워크를 소개합니다. 이 프레임워크는 두 가지 역할을 수행하는 시스템으로 구성됩니다. 하나는 비정형의 자동으로 수집된 문서 풀에서 질문과 정답을 생성하는 생성기(Generator)이고, 다른 하나는 이러한 데이터로 학습하여 성능을 향상시키는 솔버(Solver)입니다. 솔버는 생성기가 제공하는 답변에 대한 표준적인 정확성 보상을 통해 훈련되며, 생성기는 최적화 알고리즘에 대한 이해를 바탕으로 각 질문이 목표 분포에서 솔버의 성능을 실제로 향상시킬지 여부를 측정하는 영향력 점수를 통해 보상을 받습니다. 이러한 연속적이고 노이즈가 많은 영향력 점수는 표준 GRPO 방법으로는 제대로 처리하기 어렵기 때문에, 생성기 훈련을 위해 GRPO의 변형인 DuGRPO를 제안합니다. 이 모든 요소들이 함께 작용하여 문서 풀을 현재 솔버에게 유용한 질문을 우선시하는 적응형 교육 과정으로 만듭니다. Qwen3-8B-Base 모델에서 INFUSER는 강력한 자기 진화 기준 모델보다 올림피아드(Olympiad) 및 SuperGPQA 벤치마크에서 20% 이상의 상대적 성능 향상을 보였으며, 공동 진화하는 8B 크기의 생성기는 고정된 32B 크기의 추론 생성기보다 수학 및 코딩 문제 해결 능력이 뛰어납니다. 추가 실험을 통해 각 설계 요소가 필수적임을 확인했으며, INFUSER를 지시 사항에 맞춰 미세 조정된 모델에 적용하고 규칙 기반 강화 학습(RLVR) 데이터를 추가하는 두 가지 확장을 통해 프레임워크의 유연성과 일반화 가능성을 더욱 입증했습니다. 코드 및 관련 자료는 https://github.com/FFishy-git/INFUSER 에서 확인할 수 있습니다.

Original Abstract

Self-evolution offers a scalable path to stronger reasoning: a pretrained language model improves itself with only minimal external supervision. Yet existing methods either depend on extensively curated or teacher-generated training data, or, when the generator runs unsupervised, reward it by a difficulty heuristic that need not improve the solver. We introduce INFUSER, an iterative co-training framework with two co-evolving roles: a Generator that drafts questions and reference golden answers from a pool of unstructured, automatically collected documents, and a Solver that improves by training on them. The solver is trained with standard correctness rewards against the generator-provided answers, while the generator is rewarded by an optimizer-aware influence score that measures whether each proposed question would actually improve the solver on the target distribution. Because this continuous, noisy influence score is poorly served by standard GRPO, we propose DuGRPO, a dual-normalized variant of GRPO, for generator training. Together, these turn the document pool into an adaptive curriculum that favors questions useful to the current solver, not just hard ones. On Qwen3-8B-Base, INFUSER outperforms strong self-evolution baselines with over 20% relative improvement on Olympiad and SuperGPQA benchmarks, and an 8B INFUSER co-evolving generator outperforms a frozen 32B thinking generator on math and coding. Ablations confirm each design choice is necessary, and two extensions, applying INFUSER to an instruction-finetuned anchor and augmenting it with rule-verifiable RLVR data, further demonstrate the flexibility and generalizability of the framework. Code is available at https://github.com/FFishy-git/INFUSER.

0 Citations
0 Influential
23 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!