2607.12892v1 Jul 14, 2026 cs.RO

UR-VC: 비지도 로봇 가치 보정 기법을 이용한 시간 기반 진행률 추정

UR-VC: Unsupervised Robotic Value Correction for Time-Derived Progress Proxies

Ping Luo
Ping Luo
Citations: 1,863
h-index: 14
Lirui Zhao
Lirui Zhao
Citations: 722
h-index: 8
Modi Shi
Modi Shi
Citations: 502
h-index: 7
Li Chen
Li Chen
Citations: 91
h-index: 5
Qi Liu
Qi Liu
Citations: 5
h-index: 2
Hongyang Li
Hongyang Li
Citations: 724
h-index: 11

최신 로봇 학습 시스템은 중간 상태를 평가하고, 정책 학습을 안내하며, 작업 완료 여부를 감지하기 위해 빽빽한 진행 또는 가치 신호에 점점 더 의존하고 있으며, 이러한 신호의 품질이 매우 중요합니다. 이러한 빽빽한 레이블은 일반적으로 대규모로 사용 가능하지 않기 때문에, 시연 데이터 내에서 정규화된 시간이 확장 가능한 대체 수단으로 자주 사용됩니다. 즉, 후속 프레임을 진행률이 높은 것으로 간주합니다. 그러나 이 시간 기반 레이블은 실제 작업 진행률의 불완전한 근사치에 불과합니다. 접촉이 많은 조작 작업에서 로봇은 진행을 보이다가 미끄러짐, 실패한 잡기 또는 부분적인 되돌림으로 인해 진행을 잃을 수 있지만, 시간 기반 레이블은 계속해서 단조적으로 증가합니다. 본 논문에서는 오프라인 학습이 필요 없고, 별도의 학습 과정 없이 시간 기반 진행률 레이블을 수정하는 방법인 Unsupervised Robotic Value Correction (UR-VC)를 소개합니다. UR-VC는 시연 데이터의 단순한 규칙성을 활용합니다. 즉, 유사한 상태는 종종 다른 에피소드에서 다양한 타임스탬프에 반복적으로 나타납니다. UR-VC는 단일 트랙터리의 타임스탬프를 신뢰하는 대신, 다른 에피소드에서 유사한 상태를 검색하고 해당 시간 기반 레이블을 집계하여 수정된 진행률 추정치를 얻습니다. UR-VC는 수동으로 생성된 진행률 레이블이나 보상 주석 또는 추가적인 가치 모델이 필요하지 않습니다. 우리는 UR-VC를 실제 양손 조작 기반의 천 펴기 및 접기 데이터에 적용하여 평가했습니다. 이 작업은 장기간 동안 진행되는 변형 객체 조작 작업이며, 중간 진행 상태를 명확하게 보여줍니다. 수정된 레이블은 정규화된 시간이 표현할 수 없는 국소적인 회귀 및 불균일한 진행 패턴을 포착하는 반면, 전체 작업 추세를 유지합니다. 또한, 우리는 수정된 신호를 사용하여 최근의 장점 기반 정책 학습에 필요한 장점 레이블을 생성했습니다. UR-VC는 동일한 데이터, 모델 및 학습 설정을 사용했을 때 실제 로봇 작업 성공률에서 긍정적인 경향을 보였습니다.

Original Abstract

Modern robot learning systems increasingly rely on dense progress or value signals to evaluate intermediate states, guide policy learning, and detect task completion, making the quality of these signals critical. Since such dense labels are rarely available at scale, normalized time within a demonstration is often used as a scalable substitute: later frames are treated as higher progress. However, this time-derived label is only a noisy proxy for physical task progress. In contact-rich manipulation, a robot may make progress and then lose it through slips, failed grasps, or partial undoing, while the time-derived label continues to increase monotonically. We introduce Unsupervised Robotic Value Correction (UR-VC), an offline, training-free method for correcting time-derived progress labels. UR-VC exploits a simple regularity in demonstration data: similar states often recur across different episodes, but at different timestamps. Instead of trusting the timestamp from a single trajectory, UR-VC retrieves similar states from other episodes and aggregates their time-derived labels to obtain a corrected progress estimate. UR-VC requires no manual progress labels, reward annotations, or additional value model. We evaluate UR-VC on real bimanual cloth flatten-and-fold data, a long-horizon deformable-object manipulation task with visible intermediate progress. The corrected labels capture local regressions and non-uniform progress that normalized time cannot represent, while preserving the overall task trend. We further use the corrected signal to construct advantage labels for VLA training, following recent advantage-conditioned policy learning. UR-VC shows a positive trend in real-robot task success under matched data, model, and training settings.

0 Citations
0 Influential
7 Altmetric
35.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!