2606.20333v1 Jun 18, 2026 cs.AI

SoftSkill: 맥락적 적응을 위한 행동 압축

SoftSkill: Behavioral Compression for Contextual Adaptation

Yuzhi Zhao
Yuzhi Zhao
Citations: 95
h-index: 6
Ziru Liu
Ziru Liu
Citations: 65
h-index: 4
Xinyu Fu
Xinyu Fu
Citations: 58
h-index: 4
Suiyun Zhang
Suiyun Zhang
Citations: 41
h-index: 3
Xijia Tao
Xijia Tao
Citations: 134
h-index: 5
Yihua Teng
Yihua Teng
Citations: 198
h-index: 4
Kecheng Chen
Kecheng Chen
Citations: 69
h-index: 4
Ruizhe Liu
Ruizhe Liu
Citations: 10
h-index: 2
Lingpeng Kong
Lingpeng Kong
Citations: 937
h-index: 11

에이전트 기술은 일반적으로 답변 정책, 증거 활용 방식 및 작업 절차를 포함하는 자연어 Markdown 파일로 제공됩니다. 이러한 파일은 읽기 쉽고 이식성이 뛰어나지만, 간접적으로 사용됩니다. 각 작업 인스턴스마다 고정된 언어 모델이 긴 텍스트 파일을 생성 시간에 필요한 행동으로 변환해야 합니다. 본 논문에서는 자연어 기술이 고정된 기본 모델을 유지하면서, 학습 가능한 '소프트 델타'를 통해 정제된 작고 연속적인 컨텍스트 객체를 초기화할 수 있는지 질문합니다. 우리는 SoftSkill이라는 방법을 제안하는데, 이는 고정된 핵심 구조를 사용하여 다음 토큰 예측을 통해 이러한 소프트 기술을 조정하고, 추론 시간에 잠재적인 행동 선행 지식으로 활용합니다. 주요 실험에서는 Qwen3.5-4B 모델에 길이 32의 SoftSkill 접두사를 추가했을 때, Skill prompting을 사용하지 않은 경우보다 SearchQA에서 8.3점, LiveMath에서 42.1점, DocVQA에서 1.3점의 성능 향상을 보였습니다. SkillOpt와 비교하여, SoftSkill은 SearchQA에서 정확도를 5.2점, LiveMath에서 12.5점 향상시키면서, 수백 또는 수천 개의 Markdown 기술 토큰을 몇 개의 가상 토큰으로 대체합니다. 또한, 에이전트 실행이라는 더 어려운 경계 사례를 연구했는데, 여기서 희소한 경로 모방은 유용한 신호를 제공하지만, 아직 장기적인 절차적 행동을 안정적으로 압축하지 못합니다. 전반적으로, 본 논문의 결과는 일부 작업 기술이 추론 시간에 다시 해석되어야 하는 추가적인 Markdown 파일로 취급되기보다는, 고정된 모델이 작업을 시작하는 방식에 대한 작고 잠재적인 제어 방식으로 처리되는 것이 더 효과적일 수 있음을 시사합니다.

Original Abstract

Agent skills are commonly deployed as natural-language Markdown files that encode answer policies, evidence-use habits, and task procedures. These files are readable and portable, but they are consumed indirectly: for each task instance, a frozen language model must translate a long textual artifact into generation-time behavior. This paper asks whether a natural-language skill can instead initialize a compact continuous context object, refined by a trainable soft delta while the base model remains frozen. We propose SoftSkill, a frozen-backbone method that tunes such soft skills with next-token prediction and deploys them as latent behavioral priors at inference time. In our main single-round setting, a length-32 SoftSkill prefix on Qwen3.5-4B improves over no-skill prompting by 8.3 points on SearchQA, 42.1 points on LiveMath, and 1.3 points on DocVQA. Relative to SkillOpt, SoftSkill improves accuracy by 5.2 points on SearchQA and 12.5 points on LiveMath, while replacing hundreds to thousands of Markdown skill tokens with a few virtual tokens. We further study agentic execution as a harder boundary case, where sparse trajectory imitation provides useful signal but does not yet robustly compress long-horizon procedural behavior. More broadly, the results suggest that some task skills are better treated not as additional Markdown to be reinterpreted at inference time, but as compact latent controls over how a frozen model enters the task.

1 Citations
0 Influential
5.5 Altmetric
28.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!