2606.06273v1 Jun 04, 2026 cs.IT

손실 없는 픽셀 단위 이미지 전송을 위한 디퓨전 언어 모델 적용

Adapting Diffusion Language Models for Lossless Pixel-Level Image Transmission

Rongpeng Li
Rongpeng Li
Citations: 5,381
h-index: 33
Yingyue Li
Yingyue Li
Citations: 61
h-index: 3
Tianqi Ren
Tianqi Ren
Citations: 9
h-index: 1
Xianfu Chen
Xianfu Chen
Citations: 11
h-index: 2
Zhifeng Zhao
Zhifeng Zhao
Citations: 14
h-index: 3

손실 없는 픽셀 단위 이미지 전송은 의미 기반 통신보다 근본적인 영역으로, 정확한 복구를 위해서는 정확한 심볼 확률 모델링과 노이즈 채널에서의 신뢰성 있는 데이터 전송이 모두 필요합니다. 본 논문에서는 디스플레이트 디퓨전 모델을 기반으로 한 분리형 송수신 코딩 프레임워크인 DDM-SSCC를 제안하여 손실 없는 이미지 전송을 구현합니다. 기존의 라스터 순서 자동 회귀 코딩과 달리, 제안하는 송신 코덱은 픽셀 토큰 복원을 위해 디퓨전 언어 모델을 활용하고, 양방향 어텐션 하에서 동기화된 역 산술 코딩을 수행하여 여러 마스크 처리된 토큰을 하나의 역 노이즈 제거 단계 내에서 코딩할 수 있습니다. 이러한 점진적인 복원 과정은 또한 노이즈가 많은 전송 환경에서 더 유리한 송신 표현을 제공합니다. 왜냐하면 새로 복원된 토큰들은 이후의 노이즈 제거 단계에서 양방향 컨텍스트로 활용될 수 있기 때문입니다. 생성 지향 마스크 노이즈 제거와 손실 없는 산술 코딩 간의 격차를 해소하기 위해, 우리는 할턴 시퀀스를 기반으로 한 노이즈 제거 순서, 마스크 비율을 고려한 코사인 스케줄, 그리고 경량화된 온도 보정 모듈을 추가로 도입했습니다. 이러한 설계는 각각 공간적 커버리지를 향상시키고, 노이즈 제거 속도를 컨텍스트 신뢰도에 맞게 조정하며, 산술 코딩에서 사용되는 확률 테이블을 보정하는 역할을 합니다. CIFAR10, DIV2K-LR-X4 및 Kodak 데이터셋을 사용하여 가우시안 백색 잡음과 레일리 페이딩 채널 환경에서의 실험 결과, DDM-SSCC는 대표적인 손실 없는 및 의미 기반 통신 방식보다 더 우수한 정확한 복구 성능을 달성했으며, 제안된 노이즈 제거 순서, 스케줄 및 보정 모듈의 효과를 검증하는 분석도 수행했습니다.

Original Abstract

Lossless pixel-level image transmission is a fundamental regime beyond semantic communications, because exact recovery requires both accurate symbol probability modeling and reliable delivery over noisy channels. This paper proposes DDM-SSCC, a discrete-diffusion-model-based separate source-channel coding framework for lossless image transmission. Different from raster-order autoregressive coding, the proposed source codec adapts a diffusion language model to pixel-token restoration and performs synchronized reverse arithmetic coding under bidirectional attention, allowing multiple masked tokens to be coded within one reverse denoising step. This progressive restoration process also yields a more favorable source representation for noisy transmission, since newly restored tokens can serve as bidirectional context in subsequent denoising steps. To bridge the gap between generation-oriented masked denoising and lossless arithmetic coding, we further introduce a Halton-guided denoising order, a mask-ratio-aware cosine schedule, and a lightweight temperature calibration module. These designs respectively improve spatial coverage, adapt the denoising pace to context reliability, and calibrate the probability tables used by arithmetic coding. Experiments on CIFAR10, DIV2K-LR-X4, and Kodak over additive white Gaussian noise and Rayleigh fading channels show that DDM-SSCC achieves better exact-recovery performance than representative lossless and semantic communication baselines, while ablation studies verify the effectiveness of the proposed denoising order, schedule, and calibration modules.

0 Citations
0 Influential
16.5 Altmetric
82.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!