2606.20506v1 Jun 18, 2026 cs.CV

FreeStyle: 커뮤니티 LoRA 마이닝 기반 스타일-콘텐츠 이중 참조 이미지 생성의 자유로운 제어

FreeStyle: Free Control of Style-Content Dual-Reference Generation from Community LoRA Mining

Wei Cheng
Wei Cheng
Citations: 730
h-index: 10
Gang Yu
Gang Yu
Citations: 616
h-index: 9
Jinghong Lan
Jinghong Lan
Citations: 149
h-index: 4
Yunuo Chen
Yunuo Chen
Citations: 43
h-index: 4
Ziqi Ye
Ziqi Ye
Citations: 55
h-index: 4
Peng Xing
Peng Xing
Citations: 479
h-index: 5
Yixiao Fang
Yixiao Fang
Citations: 462
h-index: 3
Rui Wang
Rui Wang
Citations: 896
h-index: 7
Yufeng Yang
Yufeng Yang
Citations: 8
h-index: 2
Xuanyang Zhang
Xuanyang Zhang
Citations: 207
h-index: 7
Xianfang Zeng
Xianfang Zeng
Citations: 907
h-index: 13
Difan Zou
Difan Zou
Citations: 203
h-index: 6
Chi Zhang
Chi Zhang
Citations: 735
h-index: 7

스타일-콘텐츠 이중 참조 이미지는 콘텐츠 참조의 구조와 의미를 유지하면서 별도의 스타일 참조의 스타일을 반영하는 이미지를 합성하는 것을 목표로 합니다. 최근 상당한 발전이 있었지만, 모델은 콘텐츠 충실도, 스타일 일관성 및 지시사항 준수를 균형 있게 유지해야 하며, 스타일 참조로부터의 의미적 누출을 방지해야 하므로 이 설정은 여전히 어려운 과제입니다. 주요 병목 현상은 깨끗한 콘텐츠-스타일 분리 상태와 광범위한 롱테일 스타일 커버리지를 갖춘 대규모 삼중 데이터가 부족하다는 점입니다. 본 연구에서는 커뮤니티 LoRA 마이닝을 기반으로 하는 확장 가능한 이중 참조 이미지 생성 프레임워크인 FreeStyle을 제안합니다. 우리는 커뮤니티 LoRA를 스타일과 콘텐츠의 조립적 앵커로 간주하고, 여러 기본 모델에 걸쳐 대규모 스타일 참조 및 콘텐츠 참조 삼중 데이터를 구축하기 위해 엄격한 생성 및 필터링 파이프라인을 설계했습니다. 콘텐츠 누출 문제를 해결하기 위해, 우리는 단계별 분리 메커니즘을 갖춘 두 단계의 교육 과정을 채택합니다: 스타일 전송 단계에서 스타일 참조로부터의 누출을 억제하는 어텐션 레벨 풍부화 제약 조건과, 더 어려운 이중 참조 단계에서 위치-상응성을 기반으로 하는 누출을 대상으로 하는 주파수 인식 RoPE 변조 전략입니다. 또한, 우리는 스타일 참조 및 이중 참조 이미지 생성 모두를 포괄하는 벤치마크를 도입하고, 스타일 유사성, 콘텐츠 보존, 심미성, 지시사항 준수 및 누출 방지 측면에서 평가합니다. 벤치마크에는 스타일 불변 콘텐츠 정렬 점수(CAS)가 포함되어 있으며, 생성의 신뢰성과 누출 억제를 평가하기 위한 교정된 VLM 기반 거부 점수를 도입했습니다. 광범위한 실험 결과, 우리 모델은 스타일 일관성, 콘텐츠 보존 및 누출 억제 간에 강력한 균형을 달성함을 보여줍니다.

Original Abstract

Style-content dual-reference generation aims to synthesize an image that preserves the structure and semantics of a content reference while adopting the style of a separate style reference.Despite recent progress, this setting remains challenging because models must balance content fidelity, style alignment, and instruction following avoiding semantic leakage from the style reference.A key bottleneck is the lack of large-scale triplet data with clean content-style separation and broad long-tail style coverage.In this work, we propose FreeStyle, a scalable dual-reference generation framework based on community LoRA mining.We treat community LoRAs as compositional anchors for style and content, and design a rigorous generation and filtering pipeline to construct large-scale Style-Reference and Content-Reference triplets across multiple base models.To address content leakage, we adopt a two-stage curriculum with stage-specific disentanglement mechanisms: an attention-level enrichment constraint that suppresses style-reference leakage in the style-transfer stage, and a frequency-aware RoPE modulation strategy that targets positional-correspondence-based leakage in the harder dual-reference stage.We also introduce a benchmark covering both style-reference and dual-reference generation, with evaluations on style similarity, content preservation, aesthetics, instruction following, and leakage rejection. The benchmark incorporates a style-invariant Content Alignment Score (CAS) and introduces a calibrated VLM-based Rejection Score for evaluating generation reliability and leakage suppression.Extensive experiments show that our model achieves a strong balance among style alignment, content preservation, and leakage suppression.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!