비전 언어 모델이 비전 데이터의 개인 정보 식별 해제에 기여하는 방법
Vision Language Model Helps Private Information De-Identification in Vision Data
시각 언어 모델(VLM)은 뛰어난 능력으로 인해 큰 인기를 얻고 있습니다. 텍스트 기반 애플리케이션에서 프라이버시를 강화하는 다양한 방법이 존재하지만, 시각적 입력과 관련된 프라이버시 위험은 여전히 간과되는 경향이 있으며, 의료 이미지에 포함된 보호 건강 정보(PHI)가 대표적인 예입니다. 이러한 문제를 해결하기 위해, 민감한 텍스트를 정확하게 식별하고, 프라이버시 보호를 보장하도록 해당 데이터를 처리하는 두 가지 핵심 작업을 수행해야 합니다. 이 문제점을 해결하기 위해, VLM의 프라이버시 인지 능력을 향상시키는 데 설계된 엔드 투 엔드 프레임워크인 VisShield(Vision Privacy Shield)를 소개합니다. 우리 프레임워크는 다음과 같은 두 가지 주요 구성 요소로 이루어져 있습니다: 전문적인 지침 튜닝 데이터셋인 OPTIC(Optical Privacy Text Instruction Collection)과 맞춤형 학습 방법론입니다. 이 데이터셋은 VLM이 정밀한 민감한 텍스트 식별을 위해 특정 광학 문자 인식(OCR) 작업을 수행하도록 안내하는 다양한 프라이버시 관련 프롬프트를 제공합니다. 또한, 학습 전략은 VLM이 프라이버시 보호 작업에 효과적으로 적응하도록 보장합니다. 구체적으로, 우리 접근 방식은 VLM이 프라이버시와 관련된 텍스트를 인식하고 검출된 개체에 대한 정확한 경계 상자를 출력하도록 하여 민감한 정보의 효과적인 마스킹을 가능하게 합니다. 광범위한 실험 결과는 우리 프레임워크가 기존 방법보다 개인 정보를 처리하는 데 훨씬 뛰어난 성능을 보이며, 비전-언어 모델에서 프라이버시 보호 애플리케이션 개발에 기여할 수 있음을 보여줍니다. 데이터셋과 코드는 다음 위치에서 확인할 수 있습니다.
Visual Language Models (VLMs) have gained significant popularity due to their remarkable ability. While various methods exist to enhance privacy in text-based applications, privacy risks associated with visual inputs remain largely overlooked such as Protected Health Information (PHI) in medical images. To tackle this problem, two key tasks: accurately localizing sensitive text and processing it to ensure privacy protection should be performed. To address this issue, we introduce VisShield (Vision Privacy Shield), an end-to-end framework designed to enhance the privacy awareness of VLMs. Our framework consists of two key components: a specialized instruction-tuning dataset OPTIC (Optical Privacy Text Instruction Collection) and a tailored training methodology. The dataset provides diverse privacy-oriented prompts that guide VLMs to perform targeted Optical Character Recognition (OCR) for precise localization of sensitive text, while the training strategy ensures effective adaptation of VLMs to privacy-preserving tasks. Specifically, our approach ensures that VLMs recognize privacy-sensitive text and output precise bounding boxes for detected entities, allowing for effective masking of sensitive information. Extensive experiments demonstrate that our framework significantly outperforms existing approaches in handling private information, paving the way for privacy-preserving applications in vision-language models. Our dataset and code can be found here.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.