굴복했거나 설득되었거나: 비디오 거대 언어 모델에서의 시간적 샘플링 게이트가 주장하는 우위
Caved or Convinced: Temporal Sampling Gates Claim Deference in Video Large Language Models
두 사건 중 어떤 것이 먼저 발생했는지 질문받을 때, 비디오 거대 언어 모델은 두 가지 반대 방식으로 실패할 수 있습니다. 즉, 거짓 주장에 굴복하거나, 진실인 것을 거부할 수 있습니다. 기존 연구에서는 이러한 '비디오 아첨' 현상을 측정하고, 모델이 사용자를 덜 신뢰하도록 교육하여 이를 완화하는 방법을 제시합니다. 그러나 텍스트 및 이미지 모델에서 알려진 것처럼, 이 방법은 두 번째 유형의 오류를 악화시킬 수 있습니다. 비디오에서 발생하는 두 가지 실패는 기존 연구에서 하나의 원인으로 취급되는 두 가지 요인, 즉 '가용성(희소하게 샘플링된 프레임에 두 사건이 포함되어 있는지 여부)'과 '가중치 부여(샘플링된 증거가 사용자보다 얼마나 신뢰받는지 여부)' 때문입니다. 우리는 이러한 두 요인을 분리하기 위해 다음과 같은 두 가지 방법을 사용합니다. 첫째, 주장을 변경하는 프레임 보존 재정렬; 둘째, 고정된 프레임 예산 내에서 두 사건을 모두 포함하거나 놓치도록 샘플링 오프셋을 조정합니다. 만약 사건이 누락되면, 두 사건은 동일한 프레임을 나타내므로, 우리가 평가한 9개의 모델 각각이 진실인 주장과 거짓인 주장을 동일한 비율로 수용하며, 결과적으로 Youden's J 값이 0이 됩니다. 가용성은 필요하지만 충분하지 않습니다. 9개 모델 중 5개가 사건의 순서를 정확하게 인식했지만, 이들 중 4개는 여전히 거짓 주장에 굴복하여, '가중치 부여'라는 한계에 도달했습니다. 사용자가 제시한 증거를 기반으로 신뢰도를 조정할 수 없기 때문에, 우리는 모델이 기존에 학습한 사건 순서를 무효화하는 '역전 테스트(reversal test)'를 제안합니다. 이 테스트는 샘플링된 프레임을 순방향 및 역방향으로 분석하여 점수를 매기고, 그 결과를 바탕으로 답변하거나, 재샘플링하거나, 추측 대신 회피하도록 합니다. 이 테스트는 사건의 순서를 정확하게 인식하는 모델에서 92% ~ 100%의 정확도를 달성하며, 그렇지 않은 모델에서는 추측 대신 회피합니다.
When asked which of two events came first, video large language models can fail in two opposite ways: cave to a false claim, or reject a true one. Prior video sycophancy work measures only the first and mitigates it by teaching the model to trust the user less, a fix known in text and image models to worsen the second. In video, both failures come from two causes the literature treats as one: availability, whether the sparse sampled frames contain the two events, and weighting, whether that evidence is trusted over the user. We separate them with two interventions that keep the claim fixed: a frame-preserving reorder that flips the claim's truth, and a sampling-offset shift that captures or misses both events at a fixed frame budget. When the events are missed, the two twins present identical frames, so each of the nine models we evaluate accepts a true and a false claim at the same rate, making Youden's $J=0$ by construction. Availability is necessary but not sufficient. Five of the nine read the order, yet four of those five still cave to the false claim, so their deference hits a weighting ceiling. Since trust cannot be calibrated over evidence that was never sampled, we propose a reversal test that cancels the model's order prior by scoring the sampled frames forward and reversed, then answers, resamples, or abstains without reading the claim. The test raises the order accuracy to 0.92-1.00 on the models that read the order and abstains rather than guesses on those that cannot.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.