2608.05651v1 Aug 06, 2026 cs.CL

중계(Relay) 하세요, 라우팅하지 마세요: 비용 효율적인 LLM 기반 진화 시스템을 위한 적응형 모집단 핸드오프

Relay, Don't Route: Adaptive Population Handoff for Cost-Efficient LLM-Driven Evolution

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Guanzhi Deng
Guanzhi Deng
Citations: 27
h-index: 3
Junlan Feng
Junlan Feng
Citations: 34
h-index: 3
Haochen Luo
Haochen Luo
Citations: 134
h-index: 4
Sichun Luo
Sichun Luo
Citations: 9
h-index: 2
Zefa Hu
Zefa Hu
Citations: 10
h-index: 2
Qi Liu
Qi Liu
Citations: 9
h-index: 2
Yi Huang
Yi Huang
Citations: 104
h-index: 7
Haibo Wang
Haibo Wang
Citations: 57
h-index: 3

대규모 언어 모델(LLM) 기반 진화는 프로그램 검색 및 알고리즘 발견에 유망한 결과를 보여주었지만, 장기간의 진화 과정에서 강력한 모델에 의존하는 것은 비용이 많이 듭니다. 자연스러운 대안은 제한된 추론 예산 내에서 저렴하고 강력한 모델을 결합하는 것입니다. 그러나 기존 접근 방식은 일반적으로 개별적인 쿼리 또는 변이 단계 수준에서 모델을 할당하며, 진화 검색이 *상태 기반*이라는 점을 간과합니다. 즉, 생성된 각 후보는 이후의 변이를 발생시키는 모집단을 변경합니다. 저희는 LLM 기반 진화 경로를 경험적으로 분석하여 검색 진행은 초기 단계에 집중되는 경향이 있으며, 초기의 성능은 정보적이지만 노이즈가 많다는 점, 그리고 저렴한 모델이 상대적으로 낮은 비용으로 강력한 모델이 달성하는 초기 진행 상황의 상당 부분을 복구할 수 있다는 점을 발견했습니다. 이러한 연구 결과를 바탕으로 저희는 훈련 과정이 필요 없는 프레임워크인 ** extbf{ amemodel}**을 제안합니다. 이는 개별 호출 대신 진화하는 모집단에 예산을 할당하는 방식으로, 적응형 *모집단 핸드오프*를 통해 이루어집니다. 저렴한 모델은 뱅디트 스케줄러에 의해 할당된 짧은 블록 내에서 여러 경로를 탐색합니다. *릴레이 게인(Relay Gain)*은 핸드오프를 위해 구성된 작고 품질이 다양한 후보 집합의 주변 개선 효과로 정의되며, 이는 스케줄러 보상으로 사용되어 핸드오프 시점을 결정합니다. 선별된 후보들은 공유되는 강력한 모델 모집단을 초기화하여 추가적인 개선을 수행합니다. 저희는 네 가지 벤치마크와 세 가지 예산 설정을 사용하여 실험한 결과, ** amemodel**은 12가지 설정 중 11가지에서 경쟁적인 기본 모델보다 높은 평균 점수를 달성했습니다. 이러한 결과는 상태 기반 검색에서는 예산 할당이 개별 호출이 아닌 모집단 중심으로 이루어져야 함을 시사합니다.

Original Abstract

Large language model (LLM)-driven evolution has shown promise for program search and algorithm discovery, but relying on strong models throughout long evolutionary runs is costly. A natural alternative is to combine cheap and strong models under a fixed inference budget. However, existing approaches typically allocate models at the level of individual queries or mutation steps, overlooking that evolutionary search is \textit{stateful}: each generated candidate changes the population from which subsequent mutations are produced. We empirically analyze LLM-driven evolutionary trajectories and find that search progress is strongly front-loaded, early trajectory performance is informative but noisy, and cheap models recover much of the early progress achieved by strong models at lower cost. Motivated by these findings, we propose \textbf{\model}, a training-free framework that shifts budget allocation from individual calls to evolving populations through adaptive \textit{population handoff}. A cheap model explores multiple trajectories in short blocks allocated by a bandit scheduler. Relay Gain, defined as the marginal improvement of a compact, quality-diverse candidate bank constructed for handoff, serves as the scheduler reward and determines when to hand off. The curated candidates initialize a shared strong model population for refinement. Across four benchmarks and three budgets, \model achieves the highest mean score in 11 of 12 settings, outperforming competitive baselines. Our results suggest that in stateful search, budget allocation should be organized around the population, not the individual call.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!