2606.26981v1 Jun 25, 2026 cs.RO

문맥 기반 모델 예측 생성: 언어 모델에서 물리 법칙까지의 개방형 어휘 동작 합성

In-Context Model Predictive Generation: Open-Vocabulary Motion Synthesis from Language Models to Physics

Yang Liu
Yang Liu
Citations: 54
h-index: 2
Liang Lin
Liang Lin
Citations: 292
h-index: 8
Guanbin Li
Guanbin Li
Citations: 12,878
h-index: 56
Xiaomeng Fu
Xiaomeng Fu
Citations: 7
h-index: 2
Junfan Lin
Junfan Lin
Citations: 257
h-index: 7
Yaowei Wang
Yaowei Wang
Citations: 18
h-index: 2
Ziliang Chen
Ziliang Chen
Citations: 10
h-index: 2

텍스트 설명을 기반으로 인간 동작을 합성하는 것은 몰입형 디지털 애플리케이션에 필수적이지만, 기존 방법은 의미 충실성과 물리적 현실성 사이의 지속적인 균형 문제를 안고 있습니다. 대규모 언어 모델(LLM) 기반 접근 방식은 다양한 개방형 어휘 명령어를 해석하고 고수준 액션 계획을 구성할 수 있지만, 종종 물리적 제약을 위반하는 동작을 생성합니다. 물리 정보를 고려한 모델은 시뮬레이션 또는 제어를 통해 현실성을 향상시키지만, 의미 복잡성, 세밀한 명령어 및 새로운 개념에 대한 어려움을 겪습니다. 이러한 격차를 해결하기 위해, 우리는 문맥 기반 모델 예측 생성(ICMPG)이라는 프레임워크를 제안합니다. ICMPG는 언어 모델 계획과 추론 시 물리적 피드백을 통합하여 동작 합성을 모델 예측 제어(MPC)와 유사한 프로세스로 재구성합니다. 문맥 인식 동작 생성(CAMG) 모듈은 LLM을 사용하여 텍스트 명령어를 분해하고 동작 토큰에서 후보 동작 시퀀스를 생성합니다. 모델 예측 생성(MPG) 모듈은 이러한 후보들을 물리 시뮬레이션 및 의미 정렬을 통해 평가하고, 통합된 보상을 추정하며, 후속 생성 단계를 안내할 최적의 시퀀스를 선택합니다. 개방 루프 생성이 아닌 폐쇄 루프 개선을 통해 ICMPG는 특정 작업에 대한 정책 재학습 없이 입력 의미와 시뮬레이션된 물리 환경 모두에 동작을 적응시킬 수 있습니다. 표준 설정 및 제로샷 개방형 어휘 설정을 포함한 광범위한 실험 결과, ICMPG는 다양한 명령어에 대해 강력하게 일반화되며, 평가된 벤치마크에서 대표적인 기본 모델보다 더 물리적으로 타당하고 의미적으로 충실한 동작을 생성합니다. 이 프레임워크는 의미 해석과 물리 시뮬레이션을 연결하면서도 다양한 LLM 백본을 통합할 수 있도록 충분히 유연하여 더욱 다양하고 제어 가능한 텍스트 기반 동작 합성 기능을 제공합니다.

Original Abstract

Synthesizing human motion from textual descriptions is essential for immersive digital applications, yet existing methods face a persistent trade-off between semantic fidelity and physical realism. Large language model (LLM)-based approaches can interpret diverse open-vocabulary instructions and compose high-level action plans, but they often generate motions that violate physical constraints. Physics-aware models improve realism through simulation or control, but they struggle with semantic complexity, fine-grained instructions, and novel concepts. To address this gap, we propose In-Context Model Predictive Generation (ICMPG), a framework that integrates language-model planning with inference-time physical feedback. ICMPG reformulates motion synthesis as a Model Predictive Control (MPC)-like process with two modules. The Context-Aware Motion Generation (CAMG) module uses an LLM as a planner to decompose textual commands and generate candidate motion sequences from motion tokens. The Model Predictive Generation (MPG) module evaluates these candidates through physical simulation and semantic alignment, estimates a composite reward, and selects the best sequence to guide subsequent generation steps. Unlike open-loop generation, this closed-loop refinement enables ICMPG to adapt motions to both the input semantics and the simulated physical environment without task-specific policy retraining. Extensive experiments across standard and zero-shot open-vocabulary settings show that ICMPG generalizes robustly to diverse commands and produces motions that are more physically plausible and semantically faithful than representative baselines on the evaluated benchmarks. The framework bridges semantic interpretation and physical simulation while remaining flexible enough to incorporate different LLM backbones, enabling more versatile and controllable text-driven motion synthesis.

0 Citations
0 Influential
28 Altmetric
140.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!