장기 컨텍스트 길이를 활용한 확산 정책 학습 및 평가
Training and Evaluating Diffusion Policies with Long Context Lengths
이미테이션 학습은 RGB 이미지 관찰을 통해 로봇의 정교한 조작을 가능하게 합니다. 그러나 이러한 방법으로 훈련된 정책들은 일반적으로 로봇 동작에 대한 짧은 과거 이력 정보만을 사용합니다. 이러한 정책들은 기억을 필요로 하는 작업을 해결할 수 없으며, 반복적으로 실패하는 동작을 계속 수행하는 데 어려움을 겪습니다. 본 연구에서는 먼저 다양한 작업(지역적 안정성과 메모리 요구 사항이 다른 작업)과 여러 데이터 환경에서 컨텍스트 길이를 점진적으로 늘려가면서 정책의 성능을 비교 분석합니다. 현재까지 알려진 바와 달리, 본 연구는 이미지테이션 학습에서 컨텍스트 길이의 영향을 상세하게 조사한 최초의 연구입니다. 우리의 결과는 기존 주장에 도전하며, 단순히 컨텍스트 길이를 확장하는 것이 문헌에 제시된 것처럼 취약하지 않다는 것을 보여줍니다. 적절한 조건부 학습 방법과 노이즈 제거 네트워크(UNet+Cross-Attention)를 사용하면, 단일 작업 정책은 일반적인 데이터 환경에서 많은 작업에 대해 높은 성공률을 달성할 수 있습니다. 또한, 우리는 여러 컨텍스트 길이에 걸쳐 정책을 동시에 훈련하는 알고리즘을 제안하여 장기 컨텍스트 학습의 샘플 복잡도를 더욱 줄입니다. 마지막으로, 우리의 연구 결과를 활용하여 이전 연구에서 제안된 장기 컨텍스트 이미지테이션 학습 솔루션을 재평가합니다.
Imitation learning has enabled highly-dexterous robotic manipulation from RGB observations. Policies trained with these methods, however, typically condition robot actions on only a short history of observations. These policies cannot solve tasks that require memory and can get stuck repeatedly executing the same failing motions. In this work, we first benchmark policy performance as context length is incrementally increased from short to long, across a spectrum of tasks with varying local stability and memory requirements, and in multiple data regimes. To our knowledge, this is the first study to investigate context length in imitation learning at this level of detail. Our results challenge prior claims: naively scaling context length is not as brittle as advertised in literature. With an appropriate conditioning method and denoising backbone (UNet+Cross-Attention), single-task policies achieve high success rates on many tasks in the usual data regime even with naive scaling. Next, we propose a training algorithm to jointly train policies at multiple context lengths, further reducing the sample complexity of long-context learning. Finally, we apply our findings to re-evaluate some previously proposed solutions to long-context imitation learning.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.