생성적 추천 시스템에서 텍스트-이미지 결합 방식과 의미 기반 ID의 만남: 실증적 연구
When Text-as-Vision Meets Semantic IDs in Generative Recommendation: An Empirical Study
의미 기반 ID 학습은 생성적 추천(Generative Recommendation, GR) 모델의 핵심 인터페이스로서, 사전 학습된 텍스트 인코더를 통해 항목을 부가 정보에 기반한 이산적인 식별자로 매핑합니다. 그러나 이러한 텍스트 인코더는 주로 자연어의 문법 구조에 최적화되어 있습니다. 실제 추천 데이터에서 항목 설명은 종종 기호적이고 속성 중심적이며, 숫자, 단위 및 약어를 포함합니다. 이러한 텍스트 인코더는 이러한 정보를 단편적인 토큰으로 분리하여 의미적 일관성을 약화시키고 속성 간의 관계를 왜곡할 수 있습니다. 더욱이, 다중 모드 GR 모델로 전환할 때 표준 텍스트 인코더에 의존하면 추가적인 문제가 발생합니다. 텍스트와 이미지 임베딩은 종종 일치하지 않는 기하학적 구조를 나타내어 모달 간 융합을 덜 효과적이고 불안정하게 만듭니다. 본 논문에서는 텍스트를 시각적 신호로 간주하여 의미 기반 ID 학습을 위한 표현 방식을 재검토합니다. 항목 설명을 이미지로 렌더링하고 비전 기반 OCR 모델을 사용하여 얻은 OCR 기반 텍스트 표현에 대한 체계적인 실증 연구를 수행했습니다. 네 개의 데이터셋과 두 가지 생성 모델을 사용한 실험 결과, OCR 텍스트는 단일 모드 및 다중 모드 환경 모두에서 의미 기반 ID 학습에 있어 표준 텍스트 임베딩과 일치하거나 능가하는 성능을 보였습니다. 또한, OCR 기반 의미 기반 ID는 극단적인 공간 해상도 압축에서도 견고성을 유지하는 것으로 나타나, 실제 적용에서 강력한 성능과 효율성을 보여줍니다.
Semantic ID learning is a key interface in Generative Recommendation (GR) models, mapping items to discrete identifiers grounded in side information, most commonly via a pretrained text encoder. However, these text encoders are primarily optimized for well-formed natural language. In real-world recommendation data, item descriptions are often symbolic and attribute-centric, containing numerals, units, and abbreviations. These text encoders can break these signals into fragmented tokens, weakening semantic coherence and distorting relationships among attributes. Worse still, when moving to multimodal GR, relying on standard text encoders introduces an additional obstacle: text and image embeddings often exhibit mismatched geometric structures, making cross-modal fusion less effective and less stable. In this paper, we revisit representation design for Semantic ID learning by treating text as a visual signal. We conduct a systematic empirical study of OCR-based text representations, obtained by rendering item descriptions into images and encoding them with vision-based OCR models. Experiments across four datasets and two generative backbones show that OCR-text consistently matches or surpasses standard text embeddings for Semantic ID learning in both unimodal and multimodal settings. Furthermore, we find that OCR-based Semantic IDs remain robust under extreme spatial-resolution compression, indicating strong robustness and efficiency in practical deployments.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.