공동 정책: 반응형 인간-로봇 협업을 통한 음악 공연
Co-policy: Responsive Human-Robot Co-Creation for Musical Performances
예술은 오랫동안 인간 창의성의 중요한 표현 수단으로 자리 잡았습니다. 구체화된 인공지능은 생성 모델이 단순히 디지털 콘텐츠가 아닌 물리적 행동을 통해 그러한 창의성에 참여할 수 있는 방법을 제공합니다. 로봇 음악 공동 창작에서 의미론적 음악 이해를 실시간으로 그리고 물리적으로 실행 가능한 공연과 연결하는 것은 어려운 과제입니다. 본 논문에서는 의미론적 의도 파악, 제약 조건 내에서의 음악 변형, 시각-운동 실행을 분리하는 인간-로봇 음악 공동 창작 프레임워크인 Co-policy를 제시합니다. Co-policy는 음악의 의미론적 내용을 파악하기 위해 사전 추론 기반의 의미 고정점과 미세 조정된 Qwen-vl 플래너(F-Qwen)를 사용하여 음성, 실시간 음악 요소 및 시각 정보를 구조화된 공동 창작 계획으로 변환합니다. 낮은 지연 시간 실행을 지원하기 위해 Co-policy는 가우시안 혼합 시각-운동 정책(GMP)을 도입하며, 이는 조건부 혼합 밀도 정책으로 구현되어 대상 음표와 시각적 컨텍스트를 단일 패스에서 다중 모드 로봇 행동에 매핑합니다. 기존의 로봇 재생 시스템이 사용자가 지정한 음표를 단순히 재현하는 것과는 달리, Co-policy는 음악적 제약과 물리적 제약을 모두 고려하여 상호 보완적인 음악적 응답을 생성합니다. 실제 로봇 실험, 분석 및 전문가 평가 결과는 확산 정책 기반 모델 및 개선된 모델에 비해 의도 일치성, 실행 정확도 및 응답 빈도가 향상되었으며, 이는 구체화된 인간-AI 공동 창작을 위한 물리적으로 기반한 행동 생성의 중요성을 뒷받침합니다.
Art has long stood as a pivotal expression of human creativity. Embodied artificial intelligence offers a route for generative models to participate in that creativity through physical action rather than disembodied digital content. In robotic music co-creation, it is challenging to connect semantic musical understanding with real-time and physically executable performance. We present Co-policy, a framework for human-robot musical co-creation that separates semantic intent grounding, constrained musical variation, and visuomotor execution. To ground musical semantics, Co-policy uses pre-inference semantic anchors and a fine-tuned Qwen-vl planner (F-Qwen) to transform speech, live musical seeds, and visual observations into structured co-creation plans. To support low-latency execution, Co-policy introduces a Gaussian-Mixture Visuomotor Policy (GMP), implemented as a conditional mixture-density policy that maps target notes and visual context to multimodal robot actions in a single forward pass. Unlike robotic playback systems that merely reproduce user-specified notes, Co-policy generates complementary musical responses under both musical and physical constraints. Real-robot chime experiments, ablations, and expert evaluation show improved intent alignment, execution accuracy, and response frequency over diffusion-policy and ablated baselines, supporting physically grounded action generation as a key requirement for embodied human-AI co-creation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.