A-SR: 계층적 조정을 통한 심볼릭 회귀를 위한 자기 진화형 에이전트 기반 LLM
A-SR: Self-Evolving Agentic LLMs for Symbolic Regression via Hierarchical Coordination
심볼릭 회귀는 데이터로부터 폐쇄 형식을 갖는 방정식을 발견하는 것을 목표로 하지만, 기존의 LLM 기반 방법들은 종종 단일 제안 루프에 의존하며, 이 방식은 다양한 검색 실패를 하나의 스칼라 점수와 단일 프롬프트로 압축합니다. 본 논문에서는 표현 수정 대신 역할 기반 증거 뷰를 통해 제어 단위를 이동시키는 자기 진화형 에이전트 프레임워크인 A-SR을 제안합니다. A-SR은 조정 프로토콜 간의 라우팅, 온라인 평가기-보상 정책, 그리고 상태 기반 프로세스 메모리를 통해 공식을 발견하는 과정을 조율합니다. 검색 과정에서 평가기는 신뢰성과 생산성을 특징짓고, 역할 수준의 유틸리티를 업데이트하며, 우수한 모티프, 실패 추적 및 유효성 진단 정보를 서로 다른 에이전트로 라우팅합니다. 이 프레임워크는 두 가지 시간 척도에서 자기 진화를 수행합니다. 실행 내에서는 LLM 파라미터를 업데이트하지 않고 검색 프로세스를 적응시키며, 여러 실행을 통해 기록된 경로는 역할 기반 제안 사전 지식으로 오픈 소스 LLM에 통합될 수 있습니다. LLM-SRBench의 네 가지 LSR-Synth 과학 영역에서 평균적으로 A-SR은 Llama3.1-8B를 사용하여 기준 모델보다 Acc@0.01 값을 25.79%에서 48.30%로 향상시켰습니다. 또한, A-SR-LoRA는 Qwen3-4B의 해당 결과를 24.58%에서 38.29%로 개선했습니다. 네 가지 실제 과학 발견 작업에서 A-SR은 보고된 8가지 메트릭 중 7가지에서 가장 높은 in-distribution 또는 out-of-distribution 정규화 평균 제곱 오차를 달성했습니다.
Symbolic regression aims to discover closed-form equations from data, but existing LLM-guided methods often rely on a unified proposal loop that compresses heterogeneous search failures into a scalar score and a single prompt. We propose A-SR, a self-evolving agentic framework that shifts the control unit from expression edits to role-conditioned evidence views. A-SR coordinates formula discovery through routing among coordination protocols, an online evaluator-reward role policy, and state-routed process memory. During search, evaluator feedback characterizes reliability and productivity, updates role-level utilities, and routes elite motifs, failure traces, and validity diagnostics to different agents. The framework self-evolves at two timescales: within a run, it adapts the search process without updating LLM parameters; across runs, recorded trajectories can be distilled into open-source LLMs as role-conditioned proposal priors. Averaged over the four LSR-Synth scientific domains in LLM-SRBench, A-SR improves Acc@0.01 over baselines from 25.79% to 48.30% with Llama3.1-8B, while A-SR-LoRA improves the corresponding Qwen3-4B result from 24.58% to 38.29%. On four real-world scientific discovery tasks, A-SR obtains the best in-distribution or out-of-distribution normalized mean squared error on 7 of 8 reported metrics.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.