3차원 인간-인간 상호작용 생성에서 사회 구조가 중요하다
Social Structure Matters in 3D Human-Human Interaction Generation
텍스트 기반 동작 생성 기술이 언어를 통해 현실적인 개인 동작을 합성하는 데 상당한 발전을 이루었지만, 이를 텍스트 기반의 3차원 인간-인간 상호작용(HHI)으로 확장하는 것은 여전히 어려운 과제입니다. HHI는 단계 진행, 행위자 역할 및 행위자 간 조정을 지배하는 근본적인 사회 구조를 모델링해야 하기 때문입니다. 본 논문에서는 HHI 생성을 사회 구조 모델링 및 연결 문제로 정의합니다. 모델은 먼저 상호작용이 어떻게 전개되는지, 그리고 두 행위자가 어떻게 역할을 조정하는지를 추론하고, 이 구조를 연속적이고 물리적으로 타당하며 상대방을 고려한 3차원 동작으로 구현해야 합니다. 이러한 사회 구조가 어떻게 모델링되어야 하는지에 대한 연구를 위해, 먼저 대규모 언어 모델(LLM)이 HHI 생성에 대해 얼마나 잘 수행할 수 있는지 분석합니다. 우리의 분석 결과는 LLM이 단계 분해 및 상대방을 고려한 역할을 통해 '생각'할 수는 있지만, 동적이고 물리적으로 타당하며 상호작용에 대한 인식을 갖춘 동작을 직접 '움직임'으로 생성하는 데에는 실패한다는 것을 보여줍니다. 이러한 결과를 바탕으로, 우리는 LLM과 동작 기술을 결합하는 계획자-실행자(planner-executor) 패러다임을 제안합니다: **LLM을 사용하여 생각하고, 동작 기술을 사용하여 움직인다**. LLM 기반의 계획자는 암시적인 상호작용 의미를 단계 분해, 상대방을 고려한 행위자 역할 할당 및 동작 시퀀스와의 정렬을 통해 동작에 맞춘 사회적 제어 신호로 변환합니다. 동작 실행기는 사전 훈련된 단독 동작 모델을 LoRA, 이전 단계의 자기 조건부 학습(self-conditioning) 및 자아 상대적인 상대방 조건부 학습으로 조정하여 계획된 사회 구조를 조화로운 두 사람의 동작으로 구현합니다. 종합적으로, 우리의 Solo-to-Social 프레임워크는 사회적 조직과 동작 구현을 연결하여 더 일관성 있는 단계, 역할 정렬 및 상대방을 고려한 조화를 갖춘 3차원 HHI를 생성합니다.
Although text-to-motion generation has achieved strong progress in synthesizing realistic single-person motions from language, extending it to text-driven 3D human-human interaction (HHI) remains non-trivial, as HHI requires modeling the underlying \textbf{social structure} that governs phase progression, actor roles, and inter-actor coordination. In this paper, we formulate HHI generation as a social structure modeling and grounding problem: the model must first infer how an interaction unfolds and how the two actors coordinate their roles, and then realize this structure as continuous, physically plausible, and partner-aware 3D motion. To study how such structure should be modeled, we first examine the capability boundary of large language models (LLMs) for HHI generation. Our analysis shows that LLMs can \textit{think} by recovering phase decompositions and partner-aware roles, but cannot directly \textit{move}, as they fail to generate dynamic, physically plausible, and interaction-aware motion. This motivates our planner-executor paradigm, \textbf{Think with LLM, Move with Motion Skill}. The LLM planner converts implicit interaction semantics into motion-aligned social supervision by decomposing interactions into phases, assigning partner-aware actor roles, and aligning them with motion sequence. The motion executor then grounds the planned social structure into coordinated two-person motion by adapting a pretrained solo motion model with LoRA, previous-phase self-conditioning, and ego-relative partner conditioning. Together, our Solo-to-Social framework bridges social organization and motion realization, producing 3D HHI with improved phase consistency, role alignment, and partner-aware coordination.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.