2608.03682v1 Aug 04, 2026 cs.AI

PhyAI: 엣지에서의 실시간 물리 AI, 클라우드를 이용한 확장 가능한 배포

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud

Junbo Cui
Junbo Cui
Citations: 1,820
h-index: 6
Chenghua Wang
Chenghua Wang
Citations: 84
h-index: 3
Xuanzhe Liu
Xuanzhe Liu
Citations: 1,488
h-index: 13
Hao Zhang
Hao Zhang
Citations: 489
h-index: 5
Dongqi Cai
Dongqi Cai
Citations: 59
h-index: 3
Hao Qian
Hao Qian
Citations: 0
h-index: 0
Kezhao Zhao
Kezhao Zhao
Citations: 0
h-index: 0
Rongjie Yi
Rongjie Yi
Citations: 477
h-index: 8
Weikai Xie
Weikai Xie
Citations: 14
h-index: 2
Y. Zu
Y. Zu
Citations: 0
h-index: 0
Yuan Yao
Yuan Yao
Citations: 0
h-index: 0
Yingying Qin
Yingying Qin
Citations: 0
h-index: 0
Ziqi Guo
Ziqi Guo
Citations: 0
h-index: 0
Duojin Sun
Duojin Sun
Citations: 0
h-index: 0
Yunhan Guo
Yunhan Guo
Citations: 33
h-index: 2
Yiwen Lu
Yiwen Lu
Citations: 0
h-index: 0
Longxi Gao
Longxi Gao
Citations: 122
h-index: 6
Daliang Xu
Daliang Xu
Citations: 395
h-index: 6
Tam Sikyuen
Tam Sikyuen
Citations: 0
h-index: 0
Mengwei Xu
Mengwei Xu
Citations: 1,155
h-index: 14
Tianyue Zhang
Tianyue Zhang
Citations: 0
h-index: 0
Yuxin Zheng
Yuxin Zheng
Citations: 0
h-index: 0
Jinshuo Cui
Jinshuo Cui
Citations: 0
h-index: 0
Huaiyuan Zhang
Huaiyuan Zhang
Citations: 0
h-index: 0
Ruixin Liu
Ruixin Liu
Citations: 0
h-index: 0
Shangguang Wang
Shangguang Wang
Citations: 1,115
h-index: 21

물리 AI 정책은 모델 평가, 클라우드 강화 학습 배포, 엣지 GPU 서버 및 온보드 배치를 포함하여 전체 수명 주기 동안 추론이 필요합니다. 이러한 환경들은 동일한 체크포인트와 액션 의미를 공유하지만, 종종 별도의 추론 프로그램을 사용합니다. 이를 해결하기 위해 PhyAI는 단일 런타임을 갖춘 물리 AI 추론 엔진으로, 모델 어댑터 내에서 아키텍처별 조건 설정, 솔버, 캐시 및 출력 로직을 유지하면서 그래프 실행, 커널, 메모리 관리 및 병렬 서비스를 공유합니다. 동일한 코드는 온보드, 엣지 및 클라우드 환경에서 단일 또는 다중 GPU를 사용하여 비전-언어-액션(VLA) 모델과 월드-액션 모델(WAM)을 실행할 수 있습니다. 어댑터 인터페이스를 활용하여 MiniCPM-Robot을 출시 당일에 추가했습니다. PhyAI는 pi0, pi0.5, GR00T N1.7 및 MiniCPM-Robot의 공식 구현보다 1.40배에서 4.65배 빠른 성능을 보입니다. Cosmos3-Nano-Policy-DROID에서 8개의 H20 GPU(CFG=2, TP=4) 환경에서 지연 시간을 2.46초에서 1.18초로 줄여 2.08배의 속도 향상을 달성했습니다. 특정 구성에서는 특수 런타임이 더 빠른 성능을 보이는 경우가 있지만, 저희의 목표는 모든 경우에 가장 빠른 결과를 제공하는 것이 아니라 경쟁력 있는 지연 시간을 갖춘 단일 런타임을 만드는 것입니다. 자세한 분석 결과, 다양한 모델이 서로 다른 실행 정책을 필요로 하는 이유를 보여줍니다. Hopper 시리즈 GPU에서 배치 크기가 1일 때, pi0.5 액션 전문가가 FLOP의 8.8%를 차지하지만 지연 시간의 57.2%를 차지합니다. 배치 크기가 32일 때는 이 비율이 13.5%로 감소하고 처리량은 약 100 샘플/초에 도달합니다. Cosmos3는 여전히 생성 단계가 병목 현상을 일으키며, 배치 크기가 1에서 16으로 증가할 때 처리량이 14.3%만 향상됩니다. 또한 추론 제약과 환경 제약을 구분하는 제어 시간 Roofline을 소개하며, 측정된 pi0.5의 경우 네 가지 LIBERO 스위트에서 환경 제약에 속하는 반면, Cosmos3는 추론 제약에 속합니다. 코드 및 벤치마크: https://github.com/mingti-org/phyai.

Original Abstract

Physical AI policies require inference throughout their lifecycle, including model evaluation, cloud reinforcement learning rollout, edge GPU serving, and onboard deployment. Although these settings share the same checkpoint and action semantics, they often rely on separate inference programs. To unify them, we build PhyAI, a Physical AI inference engine with a single runtime that keeps architecture-specific conditioning, solver, cache, and output logic in model adapters while sharing graph execution, kernels, memory management, and parallel services. The same codebase runs vision-language-action (VLA) models and world-action models (WAMs) on single or multiple GPUs across onboard, edge, and cloud deployments. We used the adapter interface to add MiniCPM-Robot on the day of its release. PhyAI achieves 1.40x-4.65x speedups over the official implementations of pi0, pi0.5, GR00T N1.7, and MiniCPM-Robot. On Cosmos3-Nano-Policy-DROID it reduces latency from 2.46 to 1.18 s on eight H20 GPUs (CFG=2, TP=4), a 2.08x speedup. Specialized runtimes remain faster in several configurations, so our goal is one runtime with competitive latency rather than the fastest result in every case. Detailed profiles reveal why different models need different execution policies: on a Hopper-series GPU at batch size one, the pi0.5 action expert accounts for 8.8% of FLOPs but 57.2% of latency; at batch size 32 its share drops to 13.5% and throughput reaches about 100 samples/s. Cosmos3 remains generation-dominated and gains only 14.3% throughput as batch size increases from 1 to 16. We further introduce the control-time Roofline, which distinguishes inference-bound from environment-bound control; the measured pi0.5 points on four LIBERO suites are environment-bound while Cosmos3 stays inference-bound. Code and benchmarks: https://github.com/mingti-org/phyai.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!