OrthoMotion: 기하학적 의미론 기반 직교 어텐션을 활용한 카메라 및 피사체 움직임 분리
OrthoMotion:Disentangling Camera and Subject Motion via Geometry Semantics Orthogonal Attention
제어 가능한 비디오 생성은 카메라와 피사체의 독립적인 제어를 요구하지만, 2차원 조건 설정 방식은 이 둘을 얽히게 만듭니다. 카메라와 객체에 의해 유발되는 광류는 동일한 역심도(1/Z) 스케일링을 공유하며, 이미지 정보만으로는 분리할 수 없습니다. 본 연구에서는 이러한 얽힘이 구조적인 문제가 아니라 표현적인 문제임을 먼저 증명합니다. 즉, 2차원 카메라/객체 분리는 식별 불가능한 역문제라는 것입니다. 따라서, 분리를 연산자 설계의 문제로 재정의하고, 어텐션 연산자 수준에서 해결했습니다. OrthoMotion은 카메라 움직임을 기하학적 채널로, 피사체 움직임을 의미론적 채널로 라우팅합니다. 구체적으로, 카메라 움직임은 로터리 포지션 임베딩(RoPE)의 위상 변환을 통해 표현되며, 피사체 움직임은 크로스-어텐션에서 게이티드 값 주입을 통해 표현됩니다. 이러한 하위 연산자들은 대수적으로 상호 보완적입니다 (회전 vs. 토큰에 대한 선형 작용의 이동). 따라서, 가벼운 분리 정규화 기법은 이들의 반응 공간을 직교하도록 유도하며, 두 제어 방식 간의 상호 간섭을 방지합니다. OrthoMotion은 현재까지 알려진 바로는, 설계 자체를 통해 분리를 보장하는 첫 번째 방법입니다. 본 연구는 뛰어난 카메라 및 피사체 정확도를 동시에 달성하면서 크로스토크(cross-talk)를 최소화하며, 새로운 크로스토크 오류(CTE) 지표를 사용하여 이를 정량화했습니다. 실험 결과, OrthoMotion은 기존 방식보다 크로스토크를 2.4배 이상 줄이면서도 충실도를 유지하고 다양한 구조에 적용 가능함을 확인했습니다.
Controllable video generation demands independent command of the camera and the subject, yet 2D conditioning entangles them: camera- and object-induced optical flow share the same inverse-depth (1/Z) scaling and cannot be separated from image evidence alone. We first prove that this entanglement is representational, not architectural -- the 2D camera/object split is a non-identifiable inverse problem -- and therefore reframe decoupling as a question of operator design. We resolve it at the level of the attention operator. OrthoMotion routes camera motion into a geometric channel, a norm-preserving rotation of the rotary position embedding (RoPE) phase, and subject motion into a semantic channel, a gated value injection in cross-attention. Because these sub-operators are algebraically complementary -- a rotation versus a translation of the affine action on tokens -- a lightweight decoupling regularizer provably drives their response subspaces to orthogonality, so the two controls stop interfering. To our knowledge OrthoMotion is the first method to guarantee disentanglement by construction rather than hope for it to emerge. It attains state-of-the-art camera and subject accuracy at once while minimizing cross-talk, which we quantify with a new Cross-Talk Error (CTE) metric, cutting cross-talk by more than 2.4x with no loss in fidelity and generalizing across backbones.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.