2607.25537v1 Jul 28, 2026 cs.CV

비디오 모델을 위한 시각적 프롬프트 엔지니어링

Visual prompt engineering for video models

Kevin Swersky
Kevin Swersky
Citations: 24,079
h-index: 24
Thaddaus Wiedemer
Thaddaus Wiedemer
Citations: 198
h-index: 5
Mani Malek
Mani Malek
Citations: 1,274
h-index: 6
Oyvind Tafjord
Oyvind Tafjord
Citations: 15,012
h-index: 34
N. Kalibhat
N. Kalibhat
Citations: 156
h-index: 5
Zi Wang
Zi Wang
Citations: 3,302
h-index: 4
Robert Geirhos
Robert Geirhos
Citations: 3,789
h-index: 7
Yuxuan Li
Yuxuan Li
Citations: 197
h-index: 5
Been Kim
Been Kim
Citations: 3,404
h-index: 2
P. Jaini
P. Jaini
Citations: 2,007
h-index: 17

기초 모델 시대에, 모델의 성능은 프롬프트의 품질에 따라 결정됩니다. 이러한 이유로, 프롬프트 엔지니어링은 언어 모델 성능 향상을 위한 필수적인 기술이 되었습니다. 현재 비디오 모델이 시각적 작업(예: 시각적 추론)을 위한 기초 모델로 자리 잡고 있는 상황에서, 본 연구에서는 비디오 모델 또한 시각적 프롬프트 엔지니어링, 즉 모델 성능 개선을 위해 작업 이미지를 자동으로 수정하는 기술로부터 유사한 이점을 얻는지를 탐구합니다. 예를 들어, 시각 물리학 추론 과제(예: “장애물을 통과한 후 공이 어디에 떨어지는가?”)에서, 단순한 스케치와 같은 장면을 이미지 편집 모델 호출을 통해 사실적인 이미지로 변환할 수 있습니다. 연구 결과, 시각적 프롬프트 엔지니어링(VIPE)은 다양한 과제에서 비디오 추론 성능을 향상시키는 것으로 나타났습니다. 실제로, 비디오 모델의 경우 시각적 프롬프트 엔지니어링은 기존의 텍스트 기반 프롬프트 엔지니어링 또는 테스트 시간 스케일링보다 더욱 효과적인 방법이 될 수 있습니다. 궁극적으로, 텍스트 기반 프롬프트 엔지니어링이 언어 모델 성능을 체계적으로 향상시키는 것처럼, 시각적 프롬프트 엔지니어링은 비디오 모델로부터 더 나은 시각적 추론 성능을 이끌어내는 간단하고 계산 효율적인 접근 방식이 될 수 있습니다. 프로젝트 관련 예시 영상은 다음 웹페이지에서 확인할 수 있습니다: https://visual-prompt-engineering.github.io/.

Original Abstract

In the age of foundation models, a model is only as good as its prompt. For this reason, prompt engineering has become an essential technique for improving language model performance. Since video models are currently becoming foundation models for visual tasks (e.g., visual reasoning), we here ask whether they similarly benefit from visual prompt engineering: automatically modifying the task image to improve model performance. For example, for a visual physics reasoning task ("Where does the ball land, after passing a set of obstacles?"), an abstract sketch-like scene can be turned into a photorealistic version with a simple call to an image editing model. We find that visual prompt engineering, or VIPE for short, improves video reasoning performance across tasks. In fact, for video models, visual prompt engineering can be even more effective than classic text-based prompt engineering or test-time scaling. Ultimately, just as text-based prompt engineering systematically improves language model performance, visual prompt engineering can serve as a simple, compute-efficient approach to elicit better visual reasoning performance from video models. Example videos on our project page at https://visual-prompt-engineering.github.io/.

0 Citations
0 Influential
17 Altmetric
85.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!