2604.08294v1 Apr 09, 2026 cs.CV

시각 언어 모델은 동작 품질을 평가할 수 있는가? 경험적 평가

Can Vision Language Models Judge Action Quality? An Empirical Evaluation

P. Martins
P. Martins
Citations: 490
h-index: 6
Ricardo Rei
Ricardo Rei
Citations: 430
h-index: 11
Rui Henriques
Rui Henriques
Citations: 196
h-index: 8
Miguel Monte e Freitas
Miguel Monte e Freitas
Citations: 0
h-index: 0

동작 품질 평가(AQA)는 물리 치료, 스포츠 코칭, 그리고 경쟁 심판 분야에서 광범위하게 활용됩니다. 시각 언어 모델(VLM)은 AQA에 상당한 잠재력을 가지고 있지만, 이 분야에서의 실제 성능은 아직 제대로 규명되지 않았습니다. 본 연구에서는 최첨단 VLM 모델을 다양한 활동 영역(예: 피트니스, 피겨 스케이팅, 다이빙), 작업, 표현 방식, 그리고 프롬프트 전략 전반에 걸쳐 종합적으로 평가했습니다. 기초 결과는 Gemini 3.1 Pro, Qwen3-VL 및 InternVL3.5 모델이 무작위 추측 수준을 약간 넘어서는 성능을 보일 뿐이며, 골격 정보 포함, 지시 사항 연결, 추론 구조 활용, 그리고 in-context learning과 같은 전략들이 일시적인 성능 향상을 가져오기는 하지만, 일관적으로 효과적이지는 않다는 것을 보여줍니다. 예측 분포 분석 결과, 두 가지 체계적인 편향이 발견되었습니다. 첫째, 시각적 증거와 관계없이 정확한 실행을 예측하는 경향, 둘째, 표면적인 언어적 표현에 민감하게 반응하는 경향입니다. 이러한 편향을 완화하기 위해 작업을 재구성했지만, 개선 효과는 미미했습니다. 이는 모델의 한계가 이러한 편향을 넘어 더 근본적인 문제에 있다는 것을 시사하며, 이는 미세한 동작 품질 평가에 대한 근본적인 어려움을 나타냅니다. 본 연구는 향후 VLM 기반 AQA 연구를 위한 엄격한 기준을 제시하며, 실제 환경에 적용하기 전에 해결해야 할 문제점들을 명확하게 제시합니다.

Original Abstract

Action Quality Assessment (AQA) has broad applications in physical therapy, sports coaching, and competitive judging. Although Vision Language Models (VLMs) hold considerable promise for AQA, their actual performance in this domain remains largely uncharacterised. We present a comprehensive evaluation of state-of-the-art VLMs across activity domains (e.g. fitness, figure skating, diving), tasks, representations, and prompting strategies. Baseline results reveal that Gemini 3.1 Pro, Qwen3-VL and InternVL3.5 models perform only marginally above random chance, and although strategies such as incorporation of skeleton information, grounding instructions, reasoning structures and in-context learning lead to isolated gains, none is consistently effective. Analysis of prediction distributions uncovers two systematic biases: a tendency to predict correct execution regardless of visual evidence, and a sensitivity to superficial linguistic framing. Reformulating tasks contrastively to mitigate these biases yields minimal improvement, suggesting that the models' limitations go beyond these biases, pointing to a fundamental difficulty with fine-grained movement quality assessment. Our findings establish a rigorous baseline for future VLM-based AQA research and provide an actionable outline for failure modes requiring mitigation prior to reliable real-world deployment.

1 Citations
0 Influential
5.5 Altmetric
28.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!