2607.29172v1 Jul 31, 2026 cs.RO

CLIFT: Gemini Robotics On-Device 로봇을 비침습적 폐루프 반복 미세 조정(Fine-Tuning)을 통해 인간형 로봇 전문가로 전환

CLIFT: Turning Gemini Robotics On-Device into Humanoid Specialists via Non-Invasive Closed-Loop Iterative Fine-Tuning

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Pengcheng Wang
Pengcheng Wang
Citations: 25
h-index: 2
Masayoshi Tomizuka
Masayoshi Tomizuka
Citations: 520
h-index: 8
H. Srikanth
H. Srikanth
Citations: 5
h-index: 1
N. Jew
N. Jew
Citations: 0
h-index: 0
Junli Ren
Junli Ren
Citations: 339
h-index: 7
Peng Xu
Peng Xu
Citations: 371
h-index: 3
T. Tian
T. Tian
Citations: 36
h-index: 4

로봇 기반 모델은 점점 더 강력해지고 있지만, 가장 뛰어난 성능의 모델들은 일반적으로 독점적인 데이터로 훈련되며 폐쇄 소스 형태로 제공되어, 하위 사용자가 이러한 모델들을 새로운 작업, 형태 및 배포 환경에 맞게 조정하는 능력을 제한합니다. LLM 커뮤니티에서 나타나는 경향과 마찬가지로, 폐쇄된 가중치를 가진 로봇 기반 모델을 위한 새롭게 등장하는 접근 방식은 관리형 지도 미세 조정(Supervised Fine-Tuning, SFT) API입니다. 여기서 사용자는 훈련 데이터를 제출하고 모델 가중치, 기울기 또는 훈련 내부 정보에 대한 액세스 없이 조정된 정책을 받습니다. 이러한 API는 하위 사용자에게 강력한 독점적인 기반 모델을 활용할 수 있도록 하지만, 정책 개선을 순수한 모방으로만 제한하여 강화 학습 및 내부 훈련 신호에 의존하는 기타 폐루프 방법을 배제합니다. 이러한 제한은 특히 민첩하고 접촉이 많은 인간형 조작 작업에서 두드러지며, 새로운 상태, 동작 추적 동역학, 지연 및 컨트롤러별 오류 모드로 인해 정책 출력과 실제 작동 간의 격차가 큽니다. 본 연구에서는 관리형 API 환경에서의 인간형 로봇 적응 효과를 조사하고, 폐루프 개선을 통해 정책이 작업 숙달 수준에 도달하도록 하는 방법을 연구합니다. Gemini Robotics On-Device (GROD) 플랫폼에서 구현된 실제 인간형 로봇을 사용한 관리형 API 적응에 대한 최초의 경험적 연구 중 하나를 수행했습니다. 그 결과, API를 통한 직접적인 SFT는 동일한 데모 데이터로 훈련된 선도적인 오픈 가중치 VLA 모델보다 훨씬 뛰어난 성능을 보이지만, 민첩하고 접촉이 많은 작업에서 실제 배포 수준의 숙달에는 여전히 부족합니다. 이러한 격차를 해소하기 위해, 우리는 CLIFT(Closed-Loop Iterative Fine-Tuning)를 제안합니다. CLIFT는 배포 시의 보상 피드백을 API와 호환되는 지도 데이터로 변환하여 모델 가중치, 기울기, 확률 또는 손실에 대한 액세스 없이 폐루프 정책 개선을 가능하게 합니다. 이를 통해 GROD는 '모델 상자'를 열지 않고도 두 번의 반복 과정을 거쳐 거의 완벽한 성공률을 달성합니다.

Original Abstract

While robot foundation models are growing increasingly capable, the strongest models are typically trained on proprietary data and remain closed-source, limiting downstream users' ability to adapt them to new tasks, embodiments, and deployment settings. Following the LLM community, an emerging access paradigm for closed-weight robot foundation models is the managed supervised fine-tuning (SFT) API, where users submit training data and receive a tuned policy without access to model weights, gradients, or training internals. While such APIs let downstream users leverage powerful proprietary foundation models, they restrict policy improvement to pure imitation, ruling out reinforcement learning and other closed-loop methods that rely on internal training signals. This limitation is particularly acute for agile, contact-rich humanoid manipulation, where the gap between policy outputs and deployed behavior is large due to novel states, action tracking dynamics, latency, and controller-specific failure modes. We study how effective this managed-API regime is for humanoid adaptation, and how closed-loop improvement can be realized within it to push policies toward task mastery. We conduct one of the first empirical studies of managed-API adaptation on a real humanoid, instantiated on Gemini Robotics On-Device (GROD). We find that direct SFT through the API substantially outperforms a leading open-weight VLA trained on the same demonstrations, yet still falls short of deployment-level mastery on agile, contact-rich tasks. To close this gap, we introduce CLIFT: Closed-Loop Iterative Fine-Tuning, which turns deployment-time reward feedback into API-compatible supervised data and enables closed-loop policy improvement without accessing weights, gradients, likelihoods, or losses-pushing GROD to near-perfect success after two flywheel cycles, all without "opening the model box."

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!