2607.27557v1 Jul 30, 2026 cs.CL

자기 지도 의미 확산을 통한 매개변수와 유사한 기술 학습

Training Skills Like Parameters via Self-Supervised Semantic Diffusion

Mo Li
Mo Li
Tsinghua University
Citations: 964
h-index: 6
Zixin Yin
Zixin Yin
Citations: 285
h-index: 7
Ting Cao
Ting Cao
Citations: 73
h-index: 5
Yunxin Liu
Yunxin Liu
Citations: 98
h-index: 4

대규모 언어 모델(LLM)은 뛰어난 일반적인 지시 따르기 능력을 보여주지만, 창의적인 시나리오 작성과 같이 고도로 전문화되고 개방형 영역에서는 종종 인간 전문가에 미치지 못합니다. 기존 연구는 주로 훈련 후 단계에 초점을 맞추었지만, 지도 학습 및 강화 학습은 폐쇄 소스 모델에서 제공하지 않는 가중치 접근 권한이 필요하며 상당한 컴퓨팅 자원을 요구합니다. 또한, 학습된 내용은 특정 체크포인트에 종속되어 있으며 인간이 검토할 수 없습니다. 최근의 에이전트 기반 지속 학습 연구는 이러한 격차를 해소하기 위해 외부 텍스트 기술을 축적하려는 시도를 합니다. 그러나 이러한 방법은 비용이 많이 드는 인간 전문가의 주석이나 신뢰성이 떨어지는 LLM-as-a-judge 피드백에 크게 의존합니다. 이러한 병목 현상을 해결하기 위해, 우리는 확산 모델의 손상 및 재구성 패러다임에서 영감을 받은 새로운 비지도 자기 진화 에이전트 프레임워크를 제안합니다. 명시적인 외부 점수 부여에 의존하는 대신, 기존의 고품질 인간 제작물을 활용하여 자기 지도 신호를 생성합니다. 훈련은 신경망 학습의 일반적인 루프(순방향, 손실 계산, 역방향)를 따르며, 손실 값은 에이전트의 재구성 결과와 인간이 만든 원본 콘텐츠 간의 차이를 기반으로 합니다. 업데이트되는 것은 모델 가중치가 아닌 외부 텍스트 기술 라이브러리입니다. 우리는 우리의 프레임워크를 짧은 드라마 시나리오 작성이라는 어려운 작업에 적용하여 평가했습니다. 실험 결과는 우리가 제안하는 방법이 에이전트가 자율적으로 추출하고 내재화할 수 있는 고도로 일반화 가능한 기술을 활용하여, 특정 영역에서의 생성 능력을 크게 향상시킨다는 것을 보여줍니다. 또한, 이러한 자기 대비 방식은 외부 감독 없이 에이전트가 복잡하고 고품질의 인간 제작물을 스스로 학습하도록 하는 확장 가능한 방법을 제공합니다.

Original Abstract

While Large Language Models (LLMs) demonstrate remarkable general instruction-following capabilities, they often fall short of human experts in highly specialized, open-ended domains such as creative screenwriting. Prior approaches typically adopt post-training, yet both supervised fine-tuning and reinforcement learning require weight access that closed-source frontier models do not offer, and demand heavy compute. Moreover, what is learned is tied to a single checkpoint and cannot be inspected by humans. Recent advancements in agentic continual learning instead attempt to bridge this gap by accumulating external textual skills. However, these methods heavily rely on costly human expert annotations or unreliable LLM-as-a-judge feedback for reflection. To overcome this bottleneck, we propose a novel, unsupervised self-evolving agent framework inspired by the corruption-and-reconstruction paradigm of diffusion models. Instead of relying on explicit external scoring, we leverage existing high-quality human artifacts to construct self-supervised signals. Training then follows the familiar loop of neural network training, forward, loss, and backward, with the loss coming from contrasting the agent's reconstruction against the human original. What is updated is not model weights but an external library of textual skills. We evaluate our framework on the challenging task of short drama screenwriting. Experimental results demonstrate that our method enables the agent to autonomously extract and internalize highly generalizable skills, significantly enhancing its domain-specific generation capabilities. Furthermore, this self-contrastive reflection paradigm offers a scalable pathway for agents to teach themselves the production of complex, high-quality human artifacts, without requiring external supervision.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!