2606.27268v1 Jun 25, 2026 cs.RO

E-TTS: 로봇 조작을 위한 새로운 임베디드 테스트 시간 스케일링 프레임워크

E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation

Liang Wang
Liang Wang
Citations: 133
h-index: 4
Peiyan Li
Peiyan Li
Citations: 264
h-index: 7
Wenming Ye
Wenming Ye
Citations: 2,176
h-index: 4
Tingyu Yuan
Tingyu Yuan
Citations: 9
h-index: 2
Yuan Xu
Yuan Xu
Citations: 47
h-index: 4
Xiang Wu
Xiang Wu
Citations: 60
h-index: 3
Chaoyang Zhao
Chaoyang Zhao
Citations: 4
h-index: 1
Jing Liu
Jing Liu
Citations: 56
h-index: 5
Nianfeng Liu
Nianfeng Liu
Citations: 12
h-index: 2
Yan Huang
Yan Huang
Citations: 39
h-index: 4

최근 몇몇 연구에서 임베디드 작업에 대한 테스트 시간 스케일링을 시도했지만, 여전히 두 가지 주요 과제가 해결되지 않았습니다. (1) 추론은 정책의 성능을 향상시킬 수 있지만, 그 스케일링 메커니즘은 거의 연구되지 않았습니다. (2) 임베디드 작업은 본질적으로 장기적이고 순차적이므로 과거 정보가 필수적입니다. 따라서 현재 관찰만으로 행동 스케일링을 수행하는 것은 역사적 맥락의 부족으로 인해 충분하지 않습니다. 이러한 문제점을 해결하기 위해, 우리는 E-TTS라는 모듈화되고 쉽게 통합 가능한 임베디드 테스트 시간 스케일링 프레임워크를 제안합니다. E-TTS는 시각 및 언어 검증기를 활용한 과거 정보 기반 반복 개선을 통해 로봇 조작의 추론과 행동 스케일링을 통합합니다. E-TTS는 공동 추론-행동 스케일링을 위해 쌍으로 추론-행동 샘플링 및 점수를 수행합니다. 역사 정보를 보다 효과적으로 활용하기 위해, E-TTS는 과거 맥락을 저장하는 히스토리 버퍼를 사용하며, 이는 샘플링된 후보자를 평가하기 위해 추론 및 행동 검증기에 의해 사용됩니다. 기존의 개방형 루프 TTS 방법과는 달리, E-TTS는 샘플링 과정에 피드백 생성을 도입하여 폐쇄형 반복 개선 메커니즘을 형성하고, 추론 효율성과 환경 적응력을 향상시킵니다. 각 구성 요소는 독립적이고 결합 가능한 모듈로 작동하며, 작업 요구 사항에 따라 유연하고 적응적인 구성을 가능하게 합니다. 제안하는 프레임워크의 장점을 평가하기 위해, 4가지 서로 다른 벤치마크, 6개의 환경, 3가지 임베디드 방식 및 4가지 기본 시각-언어-행동 모델을 사용하여 실험을 진행했습니다. 실험 결과는 추가적인 전문가 데이터 수집이나 재훈련 없이도 E-TTS가 성능을 꾸준히 향상시키며, 시뮬레이션에서 최대 33.14%, 실제 환경에서는 26.62%의 성능 향상을 달성한다는 것을 보여줍니다.

Original Abstract

Recently, a few works have made early attempts to study test-time scaling for embodied tasks. However, two major challenges remain unsolved: (1) reasoning can effectively improve the performance of the policy, but its scaling mechanism has seldom been studied; (2) historical information is essential, as embodied tasks are inherently long-horizon and sequential, making sole reliance on current observations for action scaling inadequate due to the lack of historical context utilization. To address these challenges, we introduce E-TTS, a modular and plug-and-play Embodied Test-Time Scaling framework that unifies reasoning and action scaling for robotic manipulation via history-aware iterative refinement with vision-language verifiers. To support joint reasoning-action scaling, E-TTS performs reasoning-action joint sampling and scoring in a pairwise manner. To better utilize historical information, E-TTS uses a history buffer to store historical context, which is then used by reasoning and action verifiers to evaluate the sampled candidates. Unlike conventional open-loop TTS methods, E-TTS introduces feedback generation into the sampling process to form a closed-loop iterative refinement mechanism, enhancing both inference efficiency and environmental adaptability. Each component functions as an independent and composable module, allowing flexible and adaptive configuration depending on task requirements. To evaluate the advantages of our framework, we conduct experiments across 4 different benchmarks, 6 environments, 3 embodiments, and 4 base vision-language-action models. The experimental results demonstrate that, without requiring additional expert data collection or retraining, E-TTS consistently improves performance, achieving up to a 33.14% increase in simulation and 26.62% in real-world scenarios.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!