SteerVTE: 스타일 및 글리프 제어를 통한 원활한 비디오 텍스트 편집
SteerVTE: Seamless Video Text Editing with Style and Glyph Control
시각적 텍스트 편집은 이미지 및 비디오 내의 텍스트를 정확하게 수정하면서 스타일 일관성과 시각적 현실감을 유지하는 것을 목표로 합니다. 이미지 분야에서는 상당한 발전이 있었지만, 비디오 텍스트 편집은 아직 탐구되지 않은 영역입니다. 이는 작은 텍스트 영역 내에서 스트로크 수준의 정밀도를 요구하는 국소적인 작업이며, 프레임 간 정확성, 시간적 일관성 및 스타일 충실도의 어려움을 가중시킵니다. 본 논문에서는 스타일 및 글리프 제어를 통해 정확한 비디오 텍스트 편집을 수행할 수 있도록 기존 모델을 조정(steer)하는 통합 프레임워크인 SteerVTE를 소개합니다. SteerVTE는 동결된 디퓨전 트랜스포머 기반으로 구축되었으며, 가벼운 텍스트 컨텍스트 어댑터를 사용하여 원래 텍스트의 시각적 속성을 캡처하는 스타일 인코더와 라인 및 문자 수준 모두에서 대상 텍스트를 인코딩하는 이중 해상도 글리프 인코더라는 두 가지 상호 보완적인 모듈을 포함합니다. 비디오 기반 모델의 본질적으로 취약한 텍스트 렌더링 우선순위를 극복하기 위해, 우리는 글리프 정보를 고려한 공간-집중 손실 함수와 이미지에서 비디오 데이터로 점진적으로 확장되는 세 단계의 교육 방법을 제안합니다. 또한, 대규모 학습을 지원하기 위해 자동 합성 파이프라인을 개발하고 다양한 장면, 폰트 및 스타일 효과를 포함하는 백만 개의 3원소 데이터셋인 SteerVTE-1M을 구축했습니다. 광범위한 실험 결과는 SteerVTE가 텍스트 정확도, 스타일 일관성 및 시간적 일관성 측면에서 기존의 비디오 편집 모델보다 훨씬 우수한 성능을 보임을 보여줍니다.
Visual text editing aims to precisely modify text in images and videos while preserving stylistic consistency and visual realism. Despite significant advances in the image domain, video text editing remains largely unexplored: it is a localized task demanding stroke-level precision within small text regions, which compounds the challenges of cross-frame accuracy, temporal coherence, and stylistic fidelity. We introduce SteerVTE, a unified framework that \underline{\textbf{steer}}s a frozen video diffusion model to perform precise \underline{\textbf{V}}ideo \underline{\textbf{T}}ext \underline{\textbf{E}}diting through style and glyph control. Built on a frozen diffusion transformer, SteerVTE attaches a lightweight text context adapter with two complementary modules: a style encoder capturing the original text's visual attributes, and dual-granularity glyph encoders encoding the target text at both the line and character levels. To overcome the inherently weak text rendering priors of video foundation models, we further propose a glyph-aware spatial-focal loss and a three-stage progressive training curriculum that scales from image to video data. To support large-scale training, we also develop an automatic synthesis pipeline and construct SteerVTE-1M, a dataset of one million triplets spanning diverse scenes, fonts, and stylistic effects. Extensive experiments demonstrate that SteerVTE substantially outperforms existing video editing baselines across text accuracy, style consistency, and temporal coherence.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.