DIRECT: 언제 어디에서 테스트 시간 연산 자원을 할당해야 하는가? - 탑재형 계획 시스템에서의 최적화
DIRECT: When and Where Should You Allocate Test-Time Compute in Embodied Planners?
비전-언어 모델(VLMs)은 점점 더 많은 탑재형 에이전트의 고수준 계획 도구로 활용되고 있으며, 성능 향상을 위해 테스트 시간 연산 자원을 늘리는 전략이 등장하고 있습니다. 하지만 저희는 이러한 방식이 지연 시간, 토큰 사용량 및 FLOPs를 증가시키면서도 다운스트림 작업에서의 성공률 향상에 일관되지 않은, 종종 미미한 효과만을 가져와 탑재형 에이전트의 활용 범위를 제한한다는 것을 확인했습니다. 본 연구에서는 테스트 시간 연산 자원을 언제 어디에 사용하는지가 실질적인 성능 향상을 이끌어내는 데 핵심적이라는 점을 강조합니다. 저희는 다중 모드 장면 정보를 활용하여 프롬프트별로 연산 자원을 할당하는 라우팅 프레임워크인 DIRECT를 소개합니다. DIRECT는 고정된 모델 선택 방식보다 더 나은 성공-비용 파레토 최적점을 제공합니다. VLABench 및 RoboMME에서의 실험 결과, 체인 오브 씽크(Chain-of-Thought) 깊이, 모델 크기, 메모리 히스토리 등 세 가지 주요 확장 축에서 테스트 시간 연산 자원이 균일한 효과를 발휘하지 않으며, 각 축별로 질적으로 구별되는 성능 향상을 가져옴을 확인했습니다. 저희는 이러한 인사이트를 DROID 환경의 물리적인 Franka 로봇 팔 시스템에서 검증했으며, 제로샷 조작 및 장기 호라이즌 체이닝 작업을 수행하는 과정에서 저희 라우터가 더 강력한 모델과 동등하거나 높은 성공률을 달성하면서도 평균 지연 시간을 최대 65%까지 줄였습니다. 결론적으로, 저희의 연구 결과는 테스트 시간 연산 자원을 무분별하게 늘리는 것이 비효율적이며, DIRECT를 통해 로봇 시스템에서 최첨단 수준의 탑재형 계획 기능을 훨씬 낮은 비용으로 구현할 수 있음을 보여줍니다. 프로젝트 페이지는 jadee-dao.github.io/direct/ 에서 확인할 수 있습니다.
Vision-Language Models (VLMs) are increasingly deployed as high-level planners for embodied agents, with an emerging strategy of scaling test-time compute to improve capability. However, we observe that doing so increases latency, token usage, and FLOPs while yielding uneven, often diminishing gains in downstream success, limiting where embodied agents can be deployed. We argue that choosing when and where to spend test-time compute is central to bringing frontier performance to the real world. We introduce DIRECT, a routing framework that uses multimodal scene context to allocate compute per prompt, improving the success--cost Pareto frontier over fixed model selection. Across three dominant scaling axes, namely chain-of-thought depth, model size, and memory history, our experiments on VLABench and RoboMME show that test-time compute is not a uniform lever: different axes yield qualitatively distinct capability gains. We validate these insights on a physical Franka arm in a DROID setup spanning zero-shot manipulation and long-horizon chaining, where our router matches or exceeds a stronger model's success rate at up to 65% lower average latency. Ultimately, our results show that naively scaling test-time compute is wasteful, and that DIRECT can provide frontier-level embodied planning in robotic systems at a fraction of the cost. Project page can be found at jadee-dao.github.io/direct/.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.