인터랙션 트레jectory 마이닝을 통한 컴퓨터 사용 에이전트용 SKILL.md 자동 생성
Automating SKILL.md Generation for Computer-Using Agents via Interaction Trajectory Mining
명시적인 기술 라이브러리는 컴퓨터 사용 에이전트를 검사하기 쉽게 만들어 주지만, 이러한 라이브러리가 상위 레벨 정책 개선에 도움이 되는 방식으로 인터랙션 데이터에서 추출될 수 있는지 여부는 불분명합니다. 본 연구에서는 GUI 트레jectory를 분할하고, 분할된 부분을 후보 기술로 클러스터링하며, 생성된 어노테이션을 기반으로 기술 인식 정책을 학습하는 세 단계 파이프라인을 통해 이 질문에 대한 답을 찾습니다. 추출된 클러스터는 원본 벤치마크에서 가독성을 보입니다. 즉, 여덟 개의 클러스터 중 다섯 개가 InteraSkill Workflows 레이블과 최소 0.95의 순수성을 갖습니다. 그러나 가독성이 반드시 전이 성능을 의미하는 것은 아닙니다. GRPO는 IW 기술-단계 정확도를 18.5%에서 20.5%로만 향상시키고, BrowseComp+에는 거의 변화가 없으며, 주요 원본 도메인 지표에서는 단순한 빈도 기반 사전 학습보다 성능이 떨어집니다. 따라서 본 연구를 진단 연구로 제시합니다. 트레jectory 마이닝은 검사 가능한 기술 구조를 드러낼 수 있지만, 현재의 경계 감지기, 순서 없는 세그먼트 표현 및 오프라인 보상 모델은 신뢰할 수 있는 교차 도메인 정책 개선을 위해서는 충분하지 않습니다.
Explicit skill libraries make computer-using agents easier to inspect, but it remains unclear whether such libraries can be mined from interaction data in a way that improves downstream policies. We study this question through a three-stage pipeline that segments GUI trajectories, clusters segments into candidate skills, and trains a skill-aware policy from the resulting annotations. The mined clusters are readable on the source benchmark: five of eight clusters have at least 0.95 purity against InteraSkill Workflows labels. However, readability does not imply transfer. GRPO improves IW skill-step accuracy only from 18.5\% to 20.5\%, leaves BrowseComp+ essentially unchanged, and underperforms trivial frequency priors on key source-domain metrics. We therefore present the method as a diagnostic study: trajectory mining can expose inspectable skill structure, but the current boundary detector, orderless segment representation, and offline reward model are insufficient for reliable cross-domain policy improvement.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.