2608.03471v1 Aug 04, 2026 cs.CV

Hi-Token: 생성적 시각적 객체 지시를 위한 계층적 좌표 토큰화

Hi-Token: Hierarchical Coordinate Tokenization for Generative Visual Grounding

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Ke Lu
Ke Lu
Citations: 608
h-index: 9
Jian Xue
Jian Xue
Citations: 3
h-index: 1
Kun Dong
Kun Dong
Citations: 23
h-index: 3
Siwen Jiao
Siwen Jiao
Citations: 110
h-index: 4
Zijin Du
Zijin Du
Citations: 135
h-index: 4
Shun Mao
Shun Mao
Citations: 0
h-index: 0

생성형 시각-언어 모델(VLMs)은 일반적으로 경계 상자 좌표를 독립적인 출력 기호로 처리하며, 이는 숫자 순서와 축의 의미론을 명시적으로 고려하지 않습니다. 본 연구에서는 이러한 표현 방식이 시각적 객체 지시에 있어 중요한 오류 요인이 된다는 점을 밝혀냈습니다. Hi-Token은 각 좌표를 백 단위, 십 단위, 일 단위 토큰으로 인코딩하여 계층 구조를 추가하고 토큰 재사용을 증가시킵니다. 이는 기존 VLM 아키텍처를 유지하면서도 더욱 효과적인 표현 방식을 제공합니다. 또한, Hi-GAR는 그룹 상대 정책 최적화(GRPO)를 위한 기하학 기반 보상을 사용하여 경계 상자 겹침 및 여러 규모에서의 좌표 정확도를 향상시킵니다. 동일한 학습 조건 하에서 수행된 비교 실험 결과, Hi-Token은 평가된 IoU 범위 전체에 걸쳐 객체 위치 파악 성능을 향상시키는 것으로 나타났습니다. Hi-GAR는 낮은 겹침 예측을 더욱 줄이며, 학습 과정에서만 사용됩니다. 세 가지 VLM 아키텍처와 RefCOCO 데이터셋에 대한 실험 결과, 모델과 벤치마크 전반에 걸쳐 일관된 성능 향상을 보였습니다. Hi-R1은 대부분의 보고된 지표에서 강력한 특수 목적 모델보다 더 높은 값을 달성했습니다. 토큰 빈도, 숫자 경계, 객체 크기 및 IoU 분포 분석을 통해 좌표 표현 방식과 보상 기반 학습이 미치는 영향을 설명합니다. 실험 결과는 구조화된 좌표 생성이 생성형 시각적 객체 지시에 효과적인 접근 방식임을 보여줍니다.

Original Abstract

Generative Vision-Language Models (VLMs) commonly treat bounding-box coordinates as independent output symbols, leaving numerical order and axis semantics implicit. We identify this representation as an important source of error in visual grounding. Hi-Token encodes each coordinate with axis-specific tokens for the hundreds, tens, and ones digits, which adds coarse-to-fine structure and increases token reuse while retaining the existing VLM architecture. Hi-GAR complements this representation with a geometry-based reward for Group Relative Policy Optimization (GRPO), using box overlap and coordinate accuracy at multiple scales. Controlled comparisons under matched training conditions show that Hi-Token improves localization throughout the evaluated IoU range. Hi-GAR further reduces low-overlap predictions and is used only during training. Experiments on three VLM backbones and the RefCOCO family show consistent gains across models and benchmarks. Hi-R1 achieves higher values than strong specialist baselines on most reported metrics. Analyses of token frequency, digit boundaries, object scale, and IoU distributions explain the effects of coordinate representation and reward training. The results show that structured coordinate generation provides an effective approach to generative visual grounding.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!