2606.18586v1 Jun 17, 2026 cs.CV

APT: 원자적 물리 변화를 이용한 인과 관계 기반 비디오-언어 이해

APT: Atomic Physical Transitions for Causal Video-Language Understanding

Haoran Lu
Haoran Lu
Citations: 223
h-index: 6
Han Liu
Han Liu
Citations: 21
h-index: 3
Fan Du
Fan Du
Citations: 7
h-index: 1
Songlin Liu
Songlin Liu
Citations: 109
h-index: 1
Shang Wu
Shang Wu
Citations: 20
h-index: 2
Zhaoran Wang
Zhaoran Wang
Citations: 6
h-index: 2
Chenwei Xu
Chenwei Xu
Citations: 35
h-index: 4
Lie Lu
Lie Lu
Citations: 4
h-index: 1
Pranav Maneriker
Pranav Maneriker
Citations: 407
h-index: 9
Manling Li
Manling Li
Citations: 78
h-index: 4

물리적인 사건은 단순히 이름만으로는 이해될 수 없으며, 이를 구성하는 인과적인 상태 변화에 의해 이해됩니다. 예를 들어, '튕김'이라는 클립 수준의 레이블이 정확할 수는 있지만, 사건이 물리적으로 타당하게 이루어지는 과정, 즉 지지력 상실 및 접촉 시작부터 반발 및 정착까지의 과정을 숨길 수 있습니다. 이러한 숨겨진 과정을 명확히 하기 위해, 우리는 원자적 물리 변화(Atomic Physical Transitions, APT)를 도입합니다. APT는 시각적인 단서와 활성적인 물리 메커니즘을 연결하고, 시간적으로 국소화된 최소 단위의 상태 변화로서 이전/이후의 동역학적 영역을 나타냅니다. APT 체인은 비디오를 단일 집계 이벤트 레이블이 아닌 순서대로 정렬된 인과 관계 변환 시퀀스로 표현합니다. 이벤트 레이블은 '무엇'이 발생했는지를 알려주는 반면, APT 체인은 '왜' 발생했는지를 설명합니다. APT를 VLMs(Video-Language Models)가 학습할 수 있도록, 우리는 인간의 주석과 시뮬레이터 기반의 정답 데이터를 혼합하여 14가지 변환 유형을 포함하는 APT 데이터를 구축했습니다. 이는 접촉, 중력, 마찰 및 회전/안정성과 관련된 27,303개의 시간 정보를 가진 1,246개의 실험 데이터로 구성됩니다. 이 데이터를 사용하여 현재의 VLMs는 변환 수준의 물리적 지식을 제대로 이해하지 못하며, 제로샷 성능은 최대 14%에 불과하고 오류는 주로 누락된 변환에서 발생합니다. APT 체인에 대한 직접적인 파인튜닝은 변환 감지를 개선하지만 이벤트 수준에서의 학습 능력을 저하시키는데, 이는 모델이 재사용 가능한 물리적 표현을 배우는 것이 아니라 특수한 답변 형식을 학습하기 때문입니다. 따라서 우리는 VLMs가 인과 관계 기반 변환을 사용하면서도 비디오 질문에 대한 답변 능력을 잃지 않도록 하는 효율적인 방법인 APT-Tune을 제안합니다. 이는 이미지 패딩 인식 지도 학습, 형식 조건부 공동 학습 및 메커니즘 기반 도메인-유형 디코딩을 결합하여 APT 학습이 형식적으로 안정적이고 물리적으로 근거하도록 합니다. Qwen3-VL-2B 모델에 11M 개의 LoRA 파라미터만 사용하여 APT-Tune은 APT 인식률을 크게 향상시키는 동시에 이벤트 수준의 비디오 전송 성능도 개선합니다. 이러한 결과는 APT가 새로운 답변 형식이 아니라, 물리적인 비디오 이해를 위한 인간 중심적인 인과 관계 기반 지도 신호라는 것을 보여줍니다.

Original Abstract

Physical events are not understood by their names alone, but by the causal state changes that compose them. A clip-level label such as "bounce" can be correct while hiding the process that makes the event physically valid, from support loss and contact onset to rebound and settling. To make this hidden process explicit, we introduce Atomic Physical Transitions (APTs): minimal, temporally localized state changes that bind a visible cue to an active physical mechanism and before/after dynamical regimes. An APT chain represents a video as an ordered causal transition sequence rather than a single aggregate event label: event labels tell what happened; APT chains explain why it happened. To make APTs learnable by VLMs, we construct mixed-source APT data from human annotations and simulator ground truth, covering 14 transition types across contact, gravity, friction, and rotation/stability, with 27,303 timed instances over 1,246 trials. Using this data, we find that current VLMs miss transition-level physics, with zero-shot recall at most 14% and errors dominated by missed transitions. Direct fine-tuning on APT chains improves transition detection but causes event-level forgetting, indicating that the model learns a specialized answer format rather than a reusable physical representation. We therefore propose APT-Tune, a parameter-efficient recipe that teaches VLMs to use causal transitions without forgetting how to answer video questions. It combines image-pad-aware supervision, format-conditional co-training, and mechanism-conditioned domain-to-type decoding to make APT learning format-robust and physically grounded. With only 11 M LoRA parameters on Qwen3-VL-2B, APT-Tune substantially improves APT recall while also improving event-level video transfer. These results show that APTs are not a new answer format, but a human-aligned causal supervision signal for physical video understanding.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!