2603.26008v1 Mar 27, 2026 cs.CV

FairLLaVA: 공정성을 고려한 효율적인 파라미터 조정: 대규모 시각-언어 어시스턴트

FairLLaVA: Fairness-Aware Parameter-Efficient Fine-Tuning for Large Vision-Language Assistants

Tianyu Luan
Tianyu Luan
Citations: 219
h-index: 10
David S. Doermann
David S. Doermann
Citations: 1,823
h-index: 4
Mahesh Bhosale
Mahesh Bhosale
Citations: 49
h-index: 4
Abdul Wasi
Abdul Wasi
Citations: 27
h-index: 3
Shantam Srivastava
Shantam Srivastava
Citations: 1
h-index: 1
Shifa Latif
Shifa Latif
Citations: 0
h-index: 0
Mingchen Gao
Mingchen Gao
Citations: 7
h-index: 2
Xuan Gong
Xuan Gong
Citations: 4
h-index: 1

이미지 기반 생성 능력은 뛰어나지만, 다중 모달 대규모 언어 모델(MLLM)은 인구 통계 그룹별로 성능 편차가 나타나 공정성 문제를 야기할 수 있습니다. 특히 안전이 중요한 임상 환경에서 이러한 차이는 불평등한 진단 보고서를 생성하고 AI 기반 의사 결정에 대한 신뢰를 저해할 위험이 있습니다. 기존 연구에서 시각 또는 언어 모델에 대한 공정성 연구가 활발히 진행되었지만, MLLM에 대한 연구는 아직 부족합니다. 이러한 편향을 해결하기 위해, 본 논문에서는 전체 성능을 저하시키지 않으면서 시각 지시 학습에서 그룹 간의 차이를 완화하는 효율적인 파라미터 조정 방법인 FairLLaVA를 제안합니다. FairLLaVA는 대상 속성 간의 상호 정보를 최소화하여 모델의 표현을 인구 통계적으로 불변하도록 규제합니다. 이 방법은 경량 플러그인으로 통합될 수 있으며, 저차원 어댑터 기반의 파라미터 조정 방식을 사용하여 효율성을 유지하고, 다양한 아키텍처에 적용 가능한 공정한 시각 지시 학습 방법을 제공합니다. 대규모 흉부 엑스레이 보고서 생성 및 피부경 시각 질의 응답 벤치마크에 대한 광범위한 실험 결과, FairLLaVA는 그룹 간의 불평등을 지속적으로 줄이면서 다양한 의료 영상 모드에서 공정성 기반의 임상 성능과 자연어 생성 품질을 향상시키는 것으로 나타났습니다. 코드 및 관련 정보는 다음 링크에서 확인할 수 있습니다: https://github.com/bhosalems/FairLLaVA.

Original Abstract

While powerful in image-conditioned generation, multimodal large language models (MLLMs) can display uneven performance across demographic groups, highlighting fairness risks. In safety-critical clinical settings, such disparities risk producing unequal diagnostic narratives and eroding trust in AI-assisted decision-making. While fairness has been studied extensively in vision-only and language-only models, its impact on MLLMs remains largely underexplored. To address these biases, we introduce FairLLaVA, a parameter-efficient fine-tuning method that mitigates group disparities in visual instruction tuning without compromising overall performance. By minimizing the mutual information between target attributes, FairLLaVA regularizes the model's representations to be demographic-invariant. The method can be incorporated as a lightweight plug-in, maintaining efficiency with low-rank adapter fine-tuning, and provides an architecture-agnostic approach to fair visual instruction following. Extensive experiments on large-scale chest radiology report generation and dermoscopy visual question answering benchmarks show that FairLLaVA consistently reduces inter-group disparities while improving both equity-scaled clinical performance and natural language generation quality across diverse medical imaging modalities. Code can be accessed at https://github.com/bhosalems/FairLLaVA.

0 Citations
0 Influential
28.4657359028 Altmetric
0.0 Score
Original PDF
1

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!