2604.02327v1 Apr 02, 2026 cs.CV

조향 가능한 시각적 표현

Steerable Visual Representations

Jona Ruthardt
Jona Ruthardt
Citations: 1
h-index: 1
Manu Gaur
Manu Gaur
Citations: 8
h-index: 2
Makarand Tapaswi
Makarand Tapaswi
Wadhwani AI
Citations: 4,823
h-index: 25
Yuki Asano
Yuki Asano
Citations: 38
h-index: 2
D. Ramanan
D. Ramanan
Citations: 102,164
h-index: 95

DINOv2 및 MAE와 같은 사전 학습된 비전 트랜스포머(ViT)는 검색, 분류 및 분할과 같은 다양한 하위 작업에 적용할 수 있는 일반적인 이미지 특징을 제공합니다. 그러나 이러한 표현은 이미지에서 가장 중요한 시각적 단서에 집중하는 경향이 있으며, 관심 있는 덜 중요한 개념으로의 지시 방법이 없습니다. 반대로, 멀티모달 LLM은 텍스트 프롬프트를 통해 조종할 수 있지만, 결과적으로 생성되는 표현은 언어 중심적이며 일반적인 시각 작업에 대한 효과가 떨어지는 경향이 있습니다. 이러한 문제를 해결하기 위해, 우리는 자연어로 조향 가능한 전역 및 지역 특징을 갖는 새로운 유형의 시각적 표현인 '조향 가능한 시각적 표현(Steerable Visual Representations)'을 소개합니다. 대부분의 비전-언어 모델(예: CLIP)은 인코딩 후에 텍스트를 시각적 특징과 결합하는 반면(지연 결합), 우리는 경량 크로스-어텐션을 통해 텍스트를 시각적 인코더의 레이어에 직접 주입합니다(조기 결합). 우리는 표현의 조향 가능성을 측정하기 위한 벤치마크를 소개하고, 우리의 조향 가능한 시각적 특징이 이미지 내의 모든 원하는 객체에 집중하면서도 기본적인 표현 품질을 유지함을 보여줍니다. 또한, 우리의 방법은 이상 탐지 및 개인화된 객체 식별 분야에서 기존 방법과 동등하거나 우수한 성능을 보이며, 데이터 분포 외부 작업에 대한 제로샷 일반화 능력을 보여줍니다.

Original Abstract

Pretrained Vision Transformers (ViTs) such as DINOv2 and MAE provide generic image features that can be applied to a variety of downstream tasks such as retrieval, classification, and segmentation. However, such representations tend to focus on the most salient visual cues in the image, with no way to direct them toward less prominent concepts of interest. In contrast, Multimodal LLMs can be guided with textual prompts, but the resulting representations tend to be language-centric and lose their effectiveness for generic visual tasks. To address this, we introduce Steerable Visual Representations, a new class of visual representations, whose global and local features can be steered with natural language. While most vision-language models (e.g., CLIP) fuse text with visual features after encoding (late fusion), we inject text directly into the layers of the visual encoder (early fusion) via lightweight cross-attention. We introduce benchmarks for measuring representational steerability, and demonstrate that our steerable visual features can focus on any desired objects in an image while preserving the underlying representation quality. Our method also matches or outperforms dedicated approaches on anomaly detection and personalized object discrimination, exhibiting zero-shot generalization to out-of-distribution tasks.

2 Citations
0 Influential
30 Altmetric
152.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!