DeepBias: 대규모 시각-언어 모델(LVLM)의 사회적 편향에 대한 적응형 심층 분석
DeepBias: Adaptive In-depth Probing of Social Biases in LVLMs
대규모 시각-언어 모델(LVLM)은 놀라운 성능을 보여주지만, 여전히 내재된 사회적 편향에 매우 취약합니다. 기존의 편향 평가 방법은 주로 정적인 데이터셋에 의존하는데, 이는 표면적인 평가만 제공하며, 고정된 테스트 케이스로는 모델의 취약점 깊이와 한계를 정확하게 측정할 수 없습니다. 본 논문에서는 신중하게 설계된 에이전트를 사용하여 LVLM의 사회적 편향을 심층적으로 분석하는 적응형 프레임워크인 DeepBias를 소개합니다. 우리의 접근 방식은 동적인 '생성-진화-탐색' 루프를 통해 작동합니다. 먼저, 생성형 ProposerAgent는 테스트 데이터를 합성하고, 목표 LVLM의 응답을 기반으로 Direct Preference Optimization (DPO)을 통해 반복적으로 업데이트되어 모델별 오류 모드를 탐색합니다. 두 번째로, 자율적인 기술 기반 DiggerAgent는 각 테스트 데이터를 여러 번 탐색하면서, 선별된 심층화 및 재작성 전략 라이브러리에서 적응적으로 선택하여 재작성합니다. 각 단계에서 이 프로세스는 모델의 이전 응답에 의해 조건부로 제어되어 점진적으로 더 깊은 수준의 편향을 드러냅니다. 또한, 우리는 DeepBiasBench라는 벤치마크를 구축하여 우리의 프레임워크를 활용했습니다. 본 벤치마크는 최첨단 LVLM 5개를 사용하여 다양한 아키텍처에서 공유되는 취약점을 포착합니다. 종합적인 실험 결과는 우리 프레임워크의 효과성을 입증하며, DeepBias가 심층적인 편향 평가를 위한 도전적인 벤치마크임을 보여줍니다. 이는 LVLM 안전성 평가에 대한 진화적 패러다임을 제시합니다.
While Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities, they remain highly susceptible to embedded social biases. Existing bias evaluation protocols predominantly rely on static datasets, which provide only a superficial assessment, as their fixed test cases cannot adaptively evolve to measure the true depth and limits of model vulnerabilities. We introduce DeepBias, an adaptive framework for the in-depth probing of social biases in LVLMs with carefully designed agents. Our approach operates through a dynamic ''generation-evolution-probing'' loop. First, a generative ProposerAgent synthesizes test data and is iteratively updated via Direct Preference Optimization (DPO) based on the target LVLM's responses, exploring model-specific failure modes. Second, an autonomous skill-driven DiggerAgent rewrites each test data across multiple probing turns, adaptively selecting from a curated skill library of deepening and rewriting strategies. At each turn, this process is conditioned on the model's previous response, enabling progressively deeper biases to be exposed. Furthermore, we build a benchmark named DeepBiasBench using our framework. By employing an ensemble of five diverse state-of-the-art LVLMs as anchors, the benchmark captures vulnerabilities shared across architectures. Comprehensive experiments demonstrate the effectiveness of our framework and show that DeepBias provides a challenging benchmark for in-depth bias evaluation, establishing an evolutionary paradigm for LVLM safety assessment.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.