2607.25687v1 Jul 28, 2026 cs.LG

희소 관측 데이터를 이용한 도시 대기질 재구성: 결정론적 학습에서 생성 모델로

From Deterministic to Generative Deep Learning for Urban Air Quality Reconstruction from Sparse Observations

Xiaoyuan Cheng
Xiaoyuan Cheng
Citations: 46
h-index: 4
Abhishek A.Sabnis
Abhishek A.Sabnis
Citations: 0
h-index: 0
Mihai Mitrea
Mihai Mitrea
Citations: 21
h-index: 3
L. Lugon
L. Lugon
Citations: 427
h-index: 11
K. Sartelet
K. Sartelet
Citations: 221
h-index: 5
M. Bocquet
M. Bocquet
Citations: 333
h-index: 8
Shupeng Zhu
Shupeng Zhu
Citations: 983
h-index: 16
Sibo Cheng
Sibo Cheng
Citations: 1,873
h-index: 24

대기 오염 평가는 공해 노출 수준을 파악하고 공중 보건 의사 결정을 지원하는 데 필수적입니다. 그러나 대기 오염 물질 간의 복잡한 상호 작용, 예측하기 어려운 기상 패턴, 제한적인 관측소 커버리지 등의 요인으로 인해 이는 매우 복잡한 과제입니다. 본 연구에서는 딥러닝 기술을 활용하여 NO2, O3, PM2.5 및 PM10의 네 가지 주요 오염 물질에 대한 희소 관측 데이터를 기반으로 빠르고 정확한 대기질 재구성을 제공합니다. 모델은 전체 영역 시뮬레이션 데이터로 학습되었으며, 파리시 내 9개에서 28개의 관측소에서 수집된 실제 관측 데이터를 사용하여 성능을 평가했습니다. 본 연구에서는 다중 오염 물질 재구성을 위한 확산 기반 생성 프레임워크를 소개하고, 그 성능을 결정론적 딥러닝 모델과 비교합니다. 노이즈가 많은 관측 데이터와 강한 공간 변동성에도 불구하고, 제안된 모델은 시뮬레이션된 검증 데이터에서 높은 구조적 유사성을 달성했으며, 실제 관측 데이터에서도 현실적인 공간 패턴을 생성했습니다 (파워 스펙트럼 분석 결과). 또한, 모델 재학습 없이 실제 관측 데이터로의 전이를 가능하게 하는 데이터 증강 방법을 도입하여 모델이 학습 기간을 넘어 일반화될 수 있도록 했습니다. 이러한 연구 결과는 머신러닝 모델이 대기 오염 재구성 작업에 신뢰성 있는 방식으로 활용될 수 있음을 보여줍니다.

Original Abstract

Full-field reconstruction of air pollution is essential for evaluating pollution exposure and supporting public health decision-making. However, the complex interactions among pollutants, hard-to-predict weather patterns, and limited monitoring station coverage make this a complex task. We apply deep learning techniques to provide fast and accurate reconstructions from sparse observations of four key pollutants: NO2, O3, PM2.5 and PM10. Models are trained on full-field simulation data and evaluated on real-world observations collected from 9 to 28 monitoring stations in the city of Paris. We introduce a diffusion-based generative framework for multi-pollutant reconstruction and benchmark its performance against deterministic deep learning models. Despite noisy observations and strong spatial variability, the models achieve high structural similarity on simulated validation data and produce realistic spatial patterns on real-world observations, as indicated by power-spectrum analysis. We introduce data augmentation methods that enable transfer to real-world observations without retraining, allowing the models to generalise beyond the training period. These findings highlight the potential of ML models for reliable real-world deployment in air pollution reconstruction tasks.

0 Citations
0 Influential
12 Altmetric
60.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!