SkillEvo: 다중 상호작용 피드백을 통한 자기 개선형 진화 경사도
SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback
현재 에이전트 기술은 대부분 수동으로 제작되거나 단일 LLM 생성 과정을 통해 만들어지며, 따라서 실제 작동 과정에서 발생하는 오류로부터 스스로 개선할 수 있는 폐쇄 루프가 존재하지 않습니다. 최근 연구에서는 이러한 폐쇄 루프를 구축하려는 시도가 있었지만, 이는 주로 단일 턴 질의응답 평가를 기반으로 피드백을 얻습니다. 이러한 접근 방식은 심각한 비대칭성을 야기합니다. 즉, 초기 단계에서 단일 상호작용만으로는 드러낼 수 있는 문제를 해결하면 진화 경사는 감소하고, 여러 차례의 상호작용을 통해 나타나는 문제점들은 여전히 감지되지 않은 채로 남아 있으며, 결과적으로 진화는 정체됩니다. 이러한 시스템의 거버넌스 또한 엔드 투 엔드 검증 점수를 기반으로 이루어지며, 이는 성능이 저하된 후보를 거부할 수는 있지만, 그 원인을 특정하거나 수정하는 기능은 없습니다. 우리는 지속적인 기술 진화를 가로막는 요인이 편집 능력이나 반복 횟수가 아니라, 평가 피드백이 신뢰할 수 있는 진화 경사를 계속 제공하는지에 달려 있다고 주장합니다. SkillEvo는 신뢰할 수 있는 피드백을 통해 진화 경사를 생성하고, 제어 가능한 거버넌스를 통해 그 방향을 제한하는 시스템입니다. 첫 번째 구성 요소는 다중 턴 사용자 시뮬레이션을 평가의 결과 지점이 아닌 피드백 생성기로 재구성합니다. 후속 질문은 문제점을 단계별로 드러내므로, 수정 과정에서 피드백을 소비하고 동시에 새로운 피드백을 생성합니다. 두 번째 구성 요소는 수동적인 스칼라 게이트를 통한 거부 방식을 독립적인 거버넌스 레이어로 대체하여, 사실 오류 및 구조적 비효율성을 적극적으로 수정함으로써, 성능 저하가 누적되더라도 진화 경사가 왜곡되지 않도록 합니다. 6가지 유형의 클라우드 서비스, 9개의 실제 기술, 그리고 98개의 참조 파일을 대상으로 실험한 결과, SkillEvo는 자기 성찰 기반 진화를 23.0점, 단일 턴 질의응답 기반 진화를 15.4점 앞서 성능을 보였습니다.
Agent Skills are today either hand-authored or produced in a single LLM generation pass, and consequently possess no closed loop through which they might improve from the interaction failures they actually cause. Recent work does close this loop, but derives its feedback from single-turn question-answering evaluation. The consequence is a sharp asymmetry: once the first round has patched the gaps that a single exchange can reveal, the evolution gradient decays, the defects that surface only across multiple turns remain invisible, and evolution stalls. Governance in these systems is likewise driven by an end-to-end verification score, a scalar gate that can reject a degraded candidate but can neither localize nor repair its structural cause. We argue that the binding constraint on sustained skill evolution is neither editing capability nor the number of iterations, but whether the evaluation feedback keeps supplying trustworthy evolution gradients. We introduce SkillEvo, in which trustworthy feedback generates the gradient and controllable governance constrains its direction. The first component recasts multi-turn user simulation from an evaluation endpoint into a feedback generator: follow-up questions expose defects layer by layer, so that every round of revision both consumes feedback and produces new feedback. The second replaces the passive rejection of a scalar gate with an independent governance layer that actively repairs factual degradation and structural bloat, preventing the gradient from drifting as degradation accumulates. Across six categories of cloud services, 9 production Skills, and 98 skill-reference files, SkillEvo surpasses self-reflection-based evolution by 23.0 points and single- turn-QA-driven evolution by 15.4 points.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.