2608.01851v1 Aug 03, 2026 cs.RO

무게(Weights)인가, 기술(Skills)인가? 로봇 학습 기법에 대한 연구: 행동 예측 무게에서부터 스스로 기술을 개발하는 로봇까지

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

Aman Chadha
Aman Chadha
Citations: 2,232
h-index: 19
Vinija Jain
Vinija Jain
Citations: 2,079
h-index: 15
Amitava Das
Amitava Das
Citations: 942
h-index: 9
Kapil Wanaskar
Kapil Wanaskar
Citations: 20
h-index: 3
Gaytri Jena
Gaytri Jena
Citations: 8
h-index: 2
Vasu Sharma
Vasu Sharma
Citations: 19
h-index: 3

로봇 학습은 크게 두 가지 방향으로 나뉘고 있습니다. 하나는 비전-언어-행동 모델과 같이 고정된 가중치 내에 역량을 내장하는 정책 기반 접근 방식이고, 다른 하나는 에이전트가 자체적으로 실행 가능한 기술을 코드로 작성하고 개선하는 방식입니다. 본 연구는 이 '무게 대 기술'이라는 축을 중심으로 로봇 학습 분야를 정리합니다. 핵심적인 분석 내용은 코드 기반 정책 방법을 자체 개선 능력의 정도에 따라 분류하는 것입니다. 여기에는 제로샷 프로그램 합성, 폐루프 자기 수리 및 지속적인 기술 기억부터 실행 피드백, 기술 기억 및 진화적 탐색이 결합된 희소 영역까지 포함됩니다. (예: ASPIRE, ENPIRE, RoboClaw)와 같은 최신 시스템만이 이 영역에 속합니다. 또한, 비지도 강화 학습을 통한 기술 발견에서부터 대규모 언어 모델 기반의 기술 라이브러리에 이르기까지 '기술'이라는 단어가 최소 다섯 가지 이상의 서로 다른 의미로 사용되며, 그중 코드를 의미하는 경우만이 그래디언트 업데이트 없이 자체 개선이 가능하다는 점을 보여줍니다. 마지막으로, 본 연구는 이러한 분류를 로봇 기술 시장과 연관 지어 분석합니다. 현재 상용 로봇 기술 마켓플레이스에서는 '원클릭' 기술을 다양한 로봇에 제공하지만, 이는 정적인 재생 기능만을 제공하며 적응성, 이식성, 출처 추적, 안전 검증, 조합 및 표준화와 관련된 해결해야 할 과제를 제시합니다. 본 연구는 의도적으로 특정 분야에 집중하여 진행되었습니다. 전체 영역을 포괄적으로 목록화하는 대신, 77개의 대표적인 시스템을 6가지 기술 유형으로 분류하고, 하나의 분류 체계와 비교 표를 사용하여 분석하며, 각 기술 유형의 자체 개선 메커니즘에 대한 정의와 함께 해당 유형이 수행할 수 없는 작업들을 명시합니다.

Original Abstract

Robot learning is splitting into two bets: policies that bake competence into frozen weights (vision-language-action, or VLA, models), and agents that write and refine their own executable skills as code. This survey organises the field around that axis of weights versus skills. Its central analytical contribution is a deep-dive that arranges code-as-policy methods by their degree of self-improvement, from zero-shot program synthesis, through closed-loop self-repair and persistent skill memory, to the sparsely populated cell in which execution feedback, skill memory, and evolutionary search combine into one open-ended loop; only a few very recent systems (for example ASPIRE, ENPIRE, and RoboClaw) occupy that cell. We map the complementary "skills" pole, from unsupervised reinforcement-learning skill discovery to large-language-model skill libraries, and show that the word "skill" is used in at least five distinct senses, of which only the code sense self-improves without gradient updates. We then connect the taxonomy to the emerging skill economy: commercial robot-skill marketplaces now distribute one-tap skills across robots but ship only static playback, which surfaces open problems of adaptation, cross-embodiment portability, provenance, safety verification, composition, and standardisation. This is a deliberately focused survey. Rather than cataloguing the field exhaustively, it examines 77 representative systems across six technique families through one taxonomy and a set of contrast tables, and it supplies operational definitions of the self-improvement mechanisms together with a statement of what each family cannot do.

0 Citations
0 Influential
9.5 Altmetric
47.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!