레거시 데이터가 언제부터 도움이 될까? 다양한 구성 환경에서의 로봇 학습에서 나타나는 전이 현상
When Does Legacy Data Start to Help? Emergent Transfer in Cross-Configuration Robot Learning
로봇 하드웨어는 시간이 지남에 따라 진화하지만, 데모 데이터는 종종 특정 센서 및 액추에이터 구성과 연결되어 있습니다. 이는 실질적이고 아직 충분히 연구되지 않은 질문을 제기합니다: 업그레이드된 로봇에게 기존 데이터가 언제부터 도움이 될까요? 본 연구에서는 전체 형태는 동일하게 유지하면서 카메라와 그리퍼를 변경한 두 세대의 휠형 휴머노이드 플랫폼에서 이 질문을 연구했습니다. 일반적으로 다양한 구성 환경의 데이터가 항상 유용하다는 가정과는 달리, 우리는 '그로킹(grokking)'과 유사한 현상을 관찰했습니다. 즉, 업그레이드된 구성이 특정 수준의 작업 능력을 갖추기 전까지 기존 데이터는 효과적이지 않으며, 그 이후 공동 학습 효율은 급격히 증가하지만 포화에 가까워지면 감소합니다. 우리는 이와 같은 작업 의존적인 변화가 '전이 임계값'에 의해 지배된다고 가정하고, 나타나는 세 단계의 패턴을 분석했습니다. 실제 로봇 조작 작업을 통해 관찰한 결과, 낮은 능력 수준에서는 측정 가능한 효과가 없고 ($10.0% ightarrow 10.0%$), 임계값을 넘어서면 급격히 성능이 향상됩니다 (꽃 삽입 시 $23.3% ightarrow 86.7%$). 높은 능력 수준에서는 효율이 감소합니다 (펜 삽입 시 $85.0% ightarrow 93.3%$). 우리는 기울기 정렬 및 잔여 정책 불확실성을 기반으로 이론적 설명을 제공하고, 새로운 하드웨어 데이터를 언제 수집해야 하고 기존 데모를 언제 재사용해야 하는지에 대한 단계별 지침을 도출했습니다. 또한, 모바일 양팔 로봇을 이용한 물주기 작업에서도 이 세 단계 패턴을 검증했으며, 결과는 우리의 예측과 일치했습니다.
Robotic hardware evolves over time, but demonstration data is often tied to a specific sensor and actuator configuration. This raises a practical and underexplored question: when does legacy data begin to benefit an upgraded robot? We study this question on a wheeled humanoid platform across two hardware generations, where both the camera and gripper are changed while the overall morphology remains fixed. Contrary to the common assumption that more cross-configuration data is always helpful, we observe a grokking-like transition: legacy data remains ineffective until the upgraded configuration acquires a minimum level of task competence, after which co-training gains rise sharply before diminishing near saturation. We hypothesize that this task-dependent transition is governed by a transfer threshold and characterize the resulting three-phase pattern. Across real-robot manipulation tasks, we observe all three phases: no measurable benefit at low competence ($10.0\% \rightarrow 10.0\%$), a sharp gain after crossing the threshold ($23.3\% \rightarrow 86.7\%$ on flower insertion), and diminishing returns at high competence ($85.0\% \rightarrow 93.3\%$ on pen insertion). We provide a theoretical account based on gradient alignment and residual policy uncertainty, and derive a phase-aware rule for deciding when to collect more new-hardware data and when to reuse legacy demonstrations. We further validate this three-phase pattern on a mobile dual-arm watering task, with results consistent with our predictions.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.