2607.25593v1 Jul 28, 2026 cs.RO

레거시 데이터가 언제부터 도움이 될까? 다양한 구성 환경에서의 로봇 학습에서 나타나는 전이 현상

When Does Legacy Data Start to Help? Emergent Transfer in Cross-Configuration Robot Learning

Yingdong Hu
Yingdong Hu
Citations: 974
h-index: 13
Yang Gao
Yang Gao
Citations: 17
h-index: 2
Tao Wang
Tao Wang
Citations: 0
h-index: 0
Hudson Hou
Hudson Hou
Citations: 0
h-index: 0
Yufeng Liu
Yufeng Liu
Citations: 14
h-index: 2
Qinghai Li
Qinghai Li
Citations: 0
h-index: 0
Yingjie Jiang
Yingjie Jiang
Citations: 0
h-index: 0
Yingzhi Wang
Yingzhi Wang
Citations: 0
h-index: 0
Cheng Ma
Cheng Ma
Citations: 0
h-index: 0
Richard Wang
Richard Wang
Citations: 0
h-index: 0

로봇 하드웨어는 시간이 지남에 따라 진화하지만, 데모 데이터는 종종 특정 센서 및 액추에이터 구성과 연결되어 있습니다. 이는 실질적이고 아직 충분히 연구되지 않은 질문을 제기합니다: 업그레이드된 로봇에게 기존 데이터가 언제부터 도움이 될까요? 본 연구에서는 전체 형태는 동일하게 유지하면서 카메라와 그리퍼를 변경한 두 세대의 휠형 휴머노이드 플랫폼에서 이 질문을 연구했습니다. 일반적으로 다양한 구성 환경의 데이터가 항상 유용하다는 가정과는 달리, 우리는 '그로킹(grokking)'과 유사한 현상을 관찰했습니다. 즉, 업그레이드된 구성이 특정 수준의 작업 능력을 갖추기 전까지 기존 데이터는 효과적이지 않으며, 그 이후 공동 학습 효율은 급격히 증가하지만 포화에 가까워지면 감소합니다. 우리는 이와 같은 작업 의존적인 변화가 '전이 임계값'에 의해 지배된다고 가정하고, 나타나는 세 단계의 패턴을 분석했습니다. 실제 로봇 조작 작업을 통해 관찰한 결과, 낮은 능력 수준에서는 측정 가능한 효과가 없고 ($10.0% ightarrow 10.0%$), 임계값을 넘어서면 급격히 성능이 향상됩니다 (꽃 삽입 시 $23.3% ightarrow 86.7%$). 높은 능력 수준에서는 효율이 감소합니다 (펜 삽입 시 $85.0% ightarrow 93.3%$). 우리는 기울기 정렬 및 잔여 정책 불확실성을 기반으로 이론적 설명을 제공하고, 새로운 하드웨어 데이터를 언제 수집해야 하고 기존 데모를 언제 재사용해야 하는지에 대한 단계별 지침을 도출했습니다. 또한, 모바일 양팔 로봇을 이용한 물주기 작업에서도 이 세 단계 패턴을 검증했으며, 결과는 우리의 예측과 일치했습니다.

Original Abstract

Robotic hardware evolves over time, but demonstration data is often tied to a specific sensor and actuator configuration. This raises a practical and underexplored question: when does legacy data begin to benefit an upgraded robot? We study this question on a wheeled humanoid platform across two hardware generations, where both the camera and gripper are changed while the overall morphology remains fixed. Contrary to the common assumption that more cross-configuration data is always helpful, we observe a grokking-like transition: legacy data remains ineffective until the upgraded configuration acquires a minimum level of task competence, after which co-training gains rise sharply before diminishing near saturation. We hypothesize that this task-dependent transition is governed by a transfer threshold and characterize the resulting three-phase pattern. Across real-robot manipulation tasks, we observe all three phases: no measurable benefit at low competence ($10.0\% \rightarrow 10.0\%$), a sharp gain after crossing the threshold ($23.3\% \rightarrow 86.7\%$ on flower insertion), and diminishing returns at high competence ($85.0\% \rightarrow 93.3\%$ on pen insertion). We provide a theoretical account based on gradient alignment and residual policy uncertainty, and derive a phase-aware rule for deciding when to collect more new-hardware data and when to reuse legacy demonstrations. We further validate this three-phase pattern on a mobile dual-arm watering task, with results consistent with our predictions.

1 Citations
0 Influential
6.5 Altmetric
33.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!