2607.28073v1 Jul 30, 2026 cs.LG

GVR-Coder: 복잡한 문서 및 회의 시나리오에서 구조화된 SVG 생성을 위한 시각적 피드백 프레임워크

GVR-Coder: A Visual-Feedback Framework for Structured SVG Generation in Complex Document and Meeting Scenarios

Wangqiu Zhou
Wangqiu Zhou
Citations: 4
h-index: 1
Yiming Xu
Yiming Xu
Citations: 3
h-index: 1
Jihua Kang
Jihua Kang
Citations: 319
h-index: 5
Chunsai Du
Chunsai Du
Citations: 265
h-index: 2
Qifan Zhang
Qifan Zhang
Citations: 63
h-index: 3
Yiting Wu
Yiting Wu
Citations: 111
h-index: 2
Tianqi Li
Tianqi Li
Citations: 0
h-index: 0
Qiudi Song
Qiudi Song
Citations: 0
h-index: 0

전문적인 환경 및 회의 검토 시나리오에서, 긴 텍스트는 높은 인지 부하를 유발합니다. 효율적인 정보 전달을 위해, 장황한 텍스트를 논리적으로 명확한 다이어그램으로 변환하는 것이 필수적입니다. Scalable Vector Graphics (SVG)는 편집 가능성과 해상도 독립성 덕분에 이러한 목적에 효과적인 표현 방식을 제공합니다. 그러나 현재 텍스트-SVG 생성 연구는 다음과 같은 세 가지 주요 문제점에 직면해 있습니다: (1) 복잡하고 논리적인 다이어그램을 위한 데이터셋의 부족; (2) 명시적인 레이아웃 사전 지식의 부재로 인한 혼란스러운 공간 배치; (3) 렌더링된 결과물을 검증하고 미적 결함을 수정하기 위한 세밀한 시각적 피드백의 부족. 이러한 문제점을 해결하기 위해, 데이터 측면에서 문서 작성 및 회의 검토 시나리오에 특화된 대규모 SVG 데이터셋인 DocMeetSVG-100K를 소개합니다. 모델 측면에서, 우리는 고품질의 논리적인 다이어그램을 장황한 전문 텍스트로부터 생성하도록 설계된 새로운 프레임워크인 GVR-Coder를 제안합니다. 특히, 우리는 커리큘럼 기반의 거부 샘플링 미세 조정을 통해 모델이 복잡한 구조를 모델링하는 능력을 점진적으로 향상시키고, 학습 과정에서 레이아웃 제약 조건 지식을 명시적으로 통합합니다. 또한, 우리는 이중 렌더링 피드백을 통한 강화 학습 메커니즘을 도입하여 보상 신호를 통해 암묵적인 피드백을 제공함으로써 구조적 복잡성과 시각적 미학을 동시에 최적화합니다. 더욱이, 우리는 생성-검증-수정 에이전트 루프를 설계하여 명시적이고 세밀한 피드백과 목표 지향적인 개선을 통해 생성 품질을 향상시킵니다. 광범위한 실험 결과는 GVR-Coder가 경쟁 모델보다 우수한 성능을 보이며, 논리적으로 일관되고 시각적으로 매력적인 다이어그램을 안정적으로 생성한다는 것을 보여줍니다. 코드 및 데이터는 https://github.com/CurryaNa/GVR-Coder 에서 확인할 수 있습니다.

Original Abstract

In demanding professional environments and meeting review scenarios, lengthy text often imposes a high cognitive load. To facilitate efficient information communication, transforming verbose text into logically clear diagrams is essential. Scalable Vector Graphics (SVG) provide an effective representation for this purpose due to their editability and resolution independence. However, current research on Text-to-SVG generation remains hindered by three major challenges: (1) the scarcity of datasets for complex, logic-rich diagrams; (2) the absence of explicit layout priors, which leads to chaotic spatial arrangements; and (3) the lack of fine-grained visual feedback to validate rendered outputs and correct aesthetic defects. To address these challenges, at the data level, we introduce DocMeetSVG-100K, a large-scale SVG dataset tailored for document authoring and meeting review scenarios. At the model level, we propose GVR-Coder, a novel framework designed to generate high-quality logical diagrams from lengthy professional texts. Specifically, we adopt a curriculum-driven rejection sampling fine-tuning to progressively enhance the model's capability in modeling complex structures, while explicitly incorporating layout constraint knowledge during training. In addition, we introduce reinforcement learning from dual rendering feedback, a mechanism that provides implicit feedback through reward signals to jointly optimize structural complexity and visual aesthetics. Furthermore, we design a generate-verify-repair agent loop, which improves generation quality through explicit, fine-grained feedback and targeted refinement. Extensive experiments demonstrate that GVR-Coder outperforms competitive baselines and reliably produces logically coherent and visually appealing diagrams. Code and data are available at https://github.com/CurryaNa/GVR-Coder.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!