2608.07014v1 Aug 07, 2026 cs.CV

안정적인 곡선, 불안정한 항목: 비디오 LLM에서의 항목 수준별 확장 이질성

Stable Curves, Unstable Items: Item-Level Scaling Heterogeneity in Video LLMs

Hao Li
Hao Li
Citations: 55
h-index: 4
Kun Zhan
Kun Zhan
Citations: 1,574
h-index: 16
Wenzhang Sun
Wenzhang Sun
Citations: 92
h-index: 6
Chunfeng Wang
Chunfeng Wang
Citations: 10
h-index: 2
Xiangchen Yin
Xiangchen Yin
Citations: 59
h-index: 5
Yujia Chen
Yujia Chen
Citations: 0
h-index: 0

집계된 확장 곡선은 비디오 LLM이 시각적 자원(visual budget)이 증가함에 따라 안정적으로 개선되거나 포화 상태에 도달한다는 것을 시사합니다. 그러나 본 연구에서는 이러한 관점이 항목 수준에서 나타나는 크고 상반되는 변화를 가릴 수 있음을 보여줍니다. 우리는 각 고정 모델-항목 쌍을 제어된 시각적 자원 하에서의 응답 경로로 표현하고, 구성 요소의 상호 보완성, 부정적인 전환, 텍스트 오버라이트 등의 지표를 도출합니다. 세 가지 아키텍처 패밀리의 다섯 개의 공개 비디오 LLM, 네 가지 객관식 벤치마크 데이터셋, 개방형 질문 응답 및 요약, 그리고 고정된 히스토리를 가진 대화 생성 작업에 대해, 단일 시각적 자원이 모든 항목에 적합하지 않음을 확인했습니다. 네 개의 모델을 비교한 객관식 질문 응답 데이터셋에서, 항목 수준의 성능 잠재력은 8.8~18.9%p 차이를 보이며, 특정 시각적 자원 수준에서는 일부 항목이 정답이지만 다른 수준에서는 오답으로 분류됩니다. 작업에 적합한 연속적인 지표에서도 객관식 질문 응답 외에도 유사한 상호 보완성이 나타납니다. MLVU 생성의 경우 토큰-F1 점수가 2.7~3.7점 차이를, AVSD 현재 대화 생성의 경우 3.8~4.8점 차이를 보이며, 이는 평균 품질이 시각적 자원 증가에 따라 개선되더라도 나타납니다. 이러한 현상은 프레임 수, 공간 해상도, 샘플링 정책, 시간-공간 할당, 그리고 독립적으로 실행되는 원시 비디오 파이프라인 및 캐시된 파이프라인에서도 지속되며, 항목별 성능과 멤버십 추적 프로토콜 선택에 영향을 미칩니다. 의도적인 샘플링 개입을 통해 29.0%의 최종적인 성능 저하를 회복할 수 있었으며, 체계적인 프레임 검사를 통해 여러 가지 반복되는 오류 경로를 식별했습니다. 우리는 각 항목의 응답 경로, 프로토콜 출처, 생성된 어노테이션 및 재현 가능한 분석 코드를 감사 자료로 공개합니다. 또한, 특정 하이퍼파라미터(128프레임)에서 달성된 정확도를 유지하면서 평균적인 공유 프레임 비용을 31.7% 줄이는 '신뢰도 캐스케이드' 방법을 통해 응답 행렬의 실제 활용 가능성을 보여줍니다.

Original Abstract

Aggregate scaling curves suggest that Video LLMs improve smoothly or saturate as visual budgets grow. We show that this view can conceal large, opposing changes at the item level. We represent each frozen model--item pair by its response trajectory under controlled visual budgets and derive matched-grid measures of configuration complementarity, harmful transitions, and text overwrite. Across five open Video LLMs from three architecture families, four multiple-choice benchmark splits, open-ended QA and summarization, and fixed-history dialogue generation, no single budget serves all items. On the four-model matched MCQA grid, item-level oracle headroom spans $8.8$--$18.9$ accuracy points and $12.5$--$25.5\%$ of items are correct at a lower budget but wrong at a higher one. Task-appropriate continuous metrics show the same complementarity beyond multiple choice: Token-F1 oracle gaps are $2.7$--$3.7$ score points on MLVU generation and $3.8$--$4.8$ points on AVSD current-turn generation, even when mean quality improves with budget. The effect persists across frame count, spatial resolution, sampling policy, temporal--spatial allocation, and independently executed raw-video and cached pipelines, with per-item rates and membership tracking protocol choices. A controlled sampling intervention recovers $29.0\%$ of terminal regressions, and a structured frame audit identifies several recurring evidence pathways. We release per-item trajectories, protocol provenance, derived annotations, and reproducible analysis code as an auditing artifact. A confidence cascade matches fixed-$128f$ accuracy while reducing average shared frame cost by $31.7\%$, illustrating one operational use of the response matrix.

0 Citations
0 Influential
8 Altmetric
40.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!