2604.15453v1 Apr 16, 2026 cs.CV

1차원 정렬된 토큰이 효율적인 테스트 시간 검색을 가능하게 함

(1D) Ordered Tokens Enable Efficient Test-Time Search

Zhitong Gao
Zhitong Gao
Citations: 104
h-index: 6
Parham Rezaei
Parham Rezaei
Citations: 26
h-index: 3
Mingqiao Ye
Mingqiao Ye
Citations: 1,146
h-index: 5
Natavsa Jovanovi'c
Natavsa Jovanovi'c
Citations: 0
h-index: 0
Jesse Allardice
Jesse Allardice
Citations: 131
h-index: 3
Afshin Dehghan
Afshin Dehghan
Citations: 729
h-index: 13
Amir Zamir
Amir Zamir
Citations: 261
h-index: 6
Roman Bachmann
Roman Bachmann
EPFL
Citations: 1,341
h-index: 11
Ouguzhan Fatih Kar
Ouguzhan Fatih Kar
Citations: 360
h-index: 5
A. Cy
A. Cy
Citations: 123
h-index: 4

토큰화는 자기 회귀(AR) 생성 모델의 핵심 구성 요소로서, 원시 데이터를 모델링하기에 더 적합한 단위로 변환합니다. 일반적으로 토큰은 이미지의 픽셀 영역이나 텍스트의 단어 조각과 같은 지역 정보를 나타내며, AR 생성은 이러한 토큰을 고정된 순서로 예측합니다. 중요한 질문은 토큰 구조가 검증기가 여러 후보 생성 결과를 탐색하고 평가하는 테스트 시간 검색을 통한 생성을 제어하는 능력에 어떤 영향을 미치느냐입니다. 이미지 생성을 실험 대상으로 삼아, 최근의 1차원 정렬된 토크나이저로, 특히 조악함에서 세밀함으로의 구조를 가진 토크나이저가 기존의 2차원 그리드 구조보다 검색에 더 적합할 것이라는 가설을 세웠습니다. 이는 조악함에서 세밀함으로의 시퀀스에서 중간 상태가 의미론적 정보를 담고 있으며, 검증기가 이를 안정적으로 평가할 수 있어 생성 과정에서 효과적인 제어가 가능하기 때문입니다. 통제된 실험을 통해, 조악함에서 세밀함으로 정렬된 토큰으로 학습된 AR 모델이 그리드 기반 모델에 비해 향상된 테스트 시간 확장성을 보이는 것을 확인했습니다. 또한, 정렬된 구조 덕분에, AR 모델을 학습하지 않고 토큰 시퀀스에 대한 순수한 테스트 시간 검색만으로 이미지-텍스트 검증기의 지도를 받아 훈련 없이 텍스트-이미지 생성이 가능함을 보여주었습니다. 더 나아가, 고전적인 검색 알고리즘(최고 N, 빔 검색, 미리보기 검색)이 다양한 토큰 구조와 어떻게 상호 작용하는지, 그리고 다양한 검증기와 AR 사전 지식의 역할을 체계적으로 연구했습니다. 우리의 결과는 토큰 구조가 추론 시간 확장성에 미치는 영향을 강조하며, AR 모델에서의 테스트 시간 확장에 대한 실질적인 지침을 제공합니다.

Original Abstract

Tokenization is a key component of autoregressive (AR) generative models, converting raw data into more manageable units for modeling. Commonly, tokens describe local information, such as regions of pixels in images or word pieces in text, and AR generation predicts these tokens in a fixed order. A worthwhile question is whether token structures affect the ability to steer the generation through test-time search, where multiple candidate generations are explored and evaluated by a verifier. Using image generation as our testbed, we hypothesize that recent 1D ordered tokenizers with coarse-to-fine structure can be more amenable to search than classical 2D grid structures. This is rooted in the fact that the intermediate states in coarse-to-fine sequences carry semantic meaning that verifiers can reliably evaluate, enabling effective steering during generation. Through controlled experiments, we find that AR models trained on coarse-to-fine ordered tokens exhibit improved test-time scaling behavior compared to grid-based counterparts. Moreover, we demonstrate that, thanks to the ordered structure, pure test-time search over token sequences (i.e., without training an AR model) can perform training-free text-to-image generation when guided by an image-text verifier. Beyond this, we systematically study how classical search algorithms (best-of-N, beam search, lookahead search) interact with different token structures, as well as the role of different verifiers and AR priors. Our results highlight the impact of token structure on inference-time scalability and provide practical guidance for test-time scaling in AR models.

1 Citations
0 Influential
6.5 Altmetric
33.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!