DeepInsight: 물리적 AI 스택 전반에 걸친 통합 평가 인프라
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack
물리적 AI 스택의 평가는 규모가 수백 배에서 수천 배까지 다양한 연산자를 포함합니다. 여기에는 단일 파운데이션 모델 디코딩 단계부터 전체 로봇 제어 시스템의 물리 시뮬레이션까지 다양한 모달리티, 보상 의미론 및 리소스 프로필이 존재합니다. 현재 이러한 광범위한 범위를 포괄하는 프레임워크는 없으며, 따라서 스택은 별개의 평가 도구를 연결하여 평가됩니다. 이들은 런타임과 점수 체계가 다르므로 각 구성 요소의 유효성은 유지되지만, 계층 간 문제점을 진단하는 데 필요한 통합적인 관점은 상실됩니다. 본 논문에서는 단일 런타임을 통해 전체 스펙트럼을 지원하는 평가 인프라 DeepInsight를 소개합니다. DeepInsight는 다양한 환경을 동일하게 만드는 대신, 세 가지 핵심 추상화 수준인 '작업', '리소스' 및 '결과'를 유지하여 각 하위 시스템에서 공유되는 고정된 요소로 구현합니다. 여기에는 단일 에피소드 드라이버, 모든 고성능 백엔드(LLM 추론 및 격리 실행 환경 포함)에서 구현되는 리소스 핸들링 프로토콜, 그리고 모든 이벤트가 기록되는 단일 추적 식별 체계가 포함됩니다. DeepInsight는 실제 로봇 시스템의 세 가지 계층에 적용되었으며, 새로운 벤치마크를 주로 설정 변경을 통해 쉽게 통합할 수 있습니다. 기존의 성숙한 오케스트레이션 도구가 존재하는 경우(예: 파운데이션 모델 계층)에는 해당 도구의 출력을 재현하고, 동일한 테스트 스위트를 단일 노드에서 더 빠르게 실행하며, 여러 노드로 거의 선형적으로 확장합니다. DeepInsight의 가장 큰 장점은 진단 기능입니다. 모든 계층이 단일 추적에 기록하기 때문에 한 계층에서 시작되어 다른 계층에서 나타나는 문제는 해당 추적 내에서 쉽게 식별할 수 있습니다. 이는 개별 구성 요소로 분리된 평가 도구를 연결하는 방식으로는 얻을 수 없는 장점입니다.
Evaluating a Physical AI stack spans operators that differ by more than three orders of magnitude -- from a single foundation-model decoding step to thousands of physics ticks of whole-body control -- varying orthogonally in modality, reward semantics, and resource profile. No existing framework spans this range, so the stack is evaluated today by stitching together separate harnesses that share neither runtime nor scoring, preserving each segment's local validity but losing the shared identity needed to diagnose cross-layer regressions. We present DeepInsight, an evaluation infrastructure that serves this full spectrum on a single runtime. Rather than homogenize the regimes, it preserves their heterogeneity behind three narrow abstractions -- task, resource, and result -- each realized as one invariant shared by every subsystem: one episode driver, one resource-handle protocol implemented by every expensive backend (LLM inference and sandboxed runtimes alike), and one trace identity scheme under which every event is written. Deployed in production across all three layers of an embodied humanoid stack, this single set of invariants onboards new benchmarks largely by configuration. Where mature peer orchestrators exist -- at the foundation-model end -- it reproduces published references and peer-framework readings within their own spread, runs the same suites faster on a single node, and scales near-linearly across nodes. Its distinctive return is diagnostic: because every layer writes into one shared trace, a regression that begins in one layer and surfaces in another stays localizable on that trace -- a cross-layer payoff no federation of per-segment harnesses can reproduce.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.