2607.28945v1 Jul 31, 2026 cs.LG

FairDiffuseVQVAE: 표 형식 데이터 확산 모델에서의 샘플링 시간 공정성 - 벡터 양자화 잠재 변수의 조건부 세밀 조정

FairDiffuseVQVAE: Sampling-Time Fairness in Tabular Diffusion via Conditional Refinement of Vector-Quantized Latents

Amir M. Rahmani
Amir M. Rahmani
Citations: 75
h-index: 4
N. Nagesh
N. Nagesh
Citations: 182
h-index: 5
Mahdi Bagheri
Mahdi Bagheri
Citations: 5
h-index: 1

합성된 표 형식 데이터는 개인 정보 보호 데이터 공유, 데이터 증강 및 하위 분류기 편향 완화를 위해 점점 더 많이 사용되고 있습니다. TabDDPM 및 TabSyn과 같은 최첨단 표 형식 확산 모델은 뛰어난 분포 충실도를 달성하지만 공정성을 위한 메커니즘을 제공하지 않습니다. 반면에 공정성을 고려한 표 형식 데이터 생성 모델(DECAF, FairTGAN, FairTabDDPM)은 학습 시 명시적인 공정성 페널티를 적용하여 상당한 비용으로 샘플 품질 또는 하위 유틸리티에 미치는 영향이 작습니다. 우리는 Fidelity와 Fairness를 분리하는 두 단계 아키텍처인 FairDiffuseVQVAE를 소개합니다. 첫 번째 단계는 행 수준 판별기를 갖춘 벡터 양자화 오토인코더(공정성 관련 항 없음)이며, 두 번째 단계는 Stage 1의 재구성 결과와 보호 속성을 분류기 자유 가이드 방법을 통해 조건부로 조정하는 DiffuseVAE 스타일의 연속적인 확산 리파이너입니다. FairDiffuseVQVAE에서 공정성은 샘플링 분포의 특성이 됩니다. 추론 시 보호 속성에 대한 균일한 샘플링은 경쟁적인 손실 항이 아닌, 설계에 의해 인구 통계적 동등성을 보장합니다. Adult, Bank 및 COMPAS 데이터 세트에서 FairDiffuseVQVAE는 가장 높은 평균 인구 통계적 동등성 비율 ($0.702$, FairTabDDPM보다 +47%)과 동일한 기회 비율 ($0.686$, +100%)을 달성합니다. 또한, 공개된 다른 방법 중 가장 낮은 평균 쌍별 상관 오류($0.034$)를 달성하며, 이러한 공정성 향상을 위해 약 15개의 AUC 점수를 희생합니다.

Original Abstract

Synthetic tabular data is increasingly used in privacy-preserving data sharing, data augmentation, and to mitigate downstream classifier bias. State-of-the-art tabular diffusion models such as TabDDPM and TabSyn achieve excellent distributional fidelity but offer no mechanism for fairness; conversely, fairness-aware tabular generators (DECAF, FairTGAN, FairTabDDPM) impose explicit fairness penalties at training time, yielding modest fairness gains at substantial cost to either sample quality or downstream utility. We introduce FairDiffuseVQVAE, a two-stage architecture that decouples fidelity from fairness: a vector-quantized autoencoder with a row-level discriminator (Stage~1, no fairness terms) is followed by a DiffuseVAE-style continuous diffusion refiner that conditions on both the Stage-1 reconstruction and the protected attribute via classifier-free guidance (Stage~2). Fairness emerges as a property of the sampling distribution -- uniform sampling of the protected attribute at inference time enforces demographic parity by construction, rather than from competing loss terms. On the Adult, Bank and COMPAS datasets, FairDiffuseVQVAE achieves the highest mean Demographic Parity Ratio ($0.702$, $+47\%$ over FairTabDDPM) and Equalized Odds Ratio ($0.686$, $+100\%$). It also attains the lowest mean pair-wise correlation error ($0.034$) of any published method, while explicitly trading $\sim$$15$ AUC points for these fairness gains.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!