2602.07069v1 Feb 05, 2026 cs.CV

실제 이미지 초해상화를 위한 양방향 보상 기반 디퓨전 모델

Bidirectional Reward-Guided Diffusion for Real-World Image Super-Resolution

Zihao Fan
Zihao Fan
Citations: 74
h-index: 5
Xin Lu
Xin Lu
Citations: 231
h-index: 9
Jie Huang
Jie Huang
Citations: 70
h-index: 3
Xueyang Fu
Xueyang Fu
Citations: 832
h-index: 18
Zhengjun Zha
Zhengjun Zha
Citations: 761
h-index: 17
Dong Liu
Dong Liu
Citations: 10
h-index: 2
Yidi Liu
Yidi Liu
Citations: 58
h-index: 2

디퓨전 기반 초해상화는 풍부한 디테일을 생성할 수 있지만, 합성 데이터로 학습된 모델은 실제 저해상도 이미지에서 데이터 분포의 차이로 인해 성능 저하를 보이는 경우가 많습니다. 본 논문에서는 Bird-SR이라는 양방향 보상 기반 디퓨전 프레임워크를 제안합니다. Bird-SR은 초해상화를 보상 피드백 학습(ReFL)을 통해 경로 수준의 선호도 최적화 문제로 정의하며, 합성 저해상도-고해상도 쌍과 실제 저해상도 이미지를 함께 활용합니다. 구조적 충실도는 ReFL에서 쉽게 영향을 받으므로, 모델은 초기 디퓨전 단계에서 합성 쌍에 직접 최적화되어 구조 보존을 용이하게 합니다. 이는 구조 수준에서 실제 저해상도 이미지와의 분포 차이를 줄이는 데에도 도움이 됩니다. 인지적 향상을 위해, 합성 및 실제 저해상도 이미지 모두에 대해 후반 샘플링 단계에서 품질 기반 보상을 적용합니다. 보상 악용을 방지하기 위해, 합성 결과에 대한 보상은 원본 이미지에 대한 상대적 이점 공간으로 정의되며, 실제 이미지 최적화는 의미론적 정렬 제약을 통해 규제됩니다. 또한, 구조적 학습과 인지적 학습의 균형을 맞추기 위해, 초기 단계에서는 구조 보존을 강조하고, 후반 디퓨전 단계에서는 점진적으로 인지적 최적화에 초점을 맞추는 동적 충실도-인지 가중 전략을 채택합니다. 실제 초해상화 벤치마크에서의 광범위한 실험 결과, Bird-SR은 구조적 일관성을 유지하면서 인지적 품질 측면에서 최첨단 방법보다 우수한 성능을 보이며, 실제 초해상화에 효과적임을 입증합니다.

Original Abstract

Diffusion-based super-resolution can synthesize rich details, but models trained on synthetic paired data often fail on real-world LR images due to distribution shifts. We propose Bird-SR, a bidirectional reward-guided diffusion framework that formulates super-resolution as trajectory-level preference optimization via reward feedback learning (ReFL), jointly leveraging synthetic LR-HR pairs and real-world LR images. For structural fidelity easily affected in ReFL, the model is directly optimized on synthetic pairs at early diffusion steps, which also facilitates structure preservation for real-world inputs under smaller distribution gap in structure levels. For perceptual enhancement, quality-guided rewards are applied at later sampling steps to both synthetic and real LR images. To mitigate reward hacking, the rewards for synthetic results are formulated in a relative advantage space bounded by their clean counterparts, while real-world optimization is regularized via a semantic alignment constraint. Furthermore, to balance structural and perceptual learning, we adopt a dynamic fidelity-perception weighting strategy that emphasizes structure preservation at early stages and progressively shifts focus toward perceptual optimization at later diffusion steps. Extensive experiments on real-world SR benchmarks demonstrate that Bird-SR consistently outperforms state-of-the-art methods in perceptual quality while preserving structural consistency, validating its effectiveness for real-world super-resolution.

1 Citations
0 Influential
9 Altmetric
46.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!