AdaCultureSafe: 대규모 언어 모델에서 문화적 지식을 기반으로 한 적응형 문화적 안전성
AdaCultureSafe: Adaptive Cultural Safety Grounded by Cultural Knowledge in Large Language Models
대규모 언어 모델(LLM)의 광범위한 사용으로 인해, 모델의 문화적 안전성과 책임감 있는 글로벌 적용을 위해서는 토착 문화를 존중하는 것이 필수적입니다. 기존 연구들은 문화적 안전성과 문화적 지식을 별도로 고려하며, 전자가 후자에 기반해야 한다는 점을 간과합니다. 이는 LLM이 문화에 특정한 존중하는 답변을 생성하는 데 심각한 제약을 초래합니다. 결과적으로, 적응형 문화적 안전성은 여전히 해결해야 할 중요한 과제입니다. 본 연구에서는 문화적 안전성과 지식을 동시에 모델링하는 방법을 제안합니다. 무엇보다 중요한 것은, 본 연구를 수행하기 위한 핵심 전제 조건은 문화적 안전성과 지식을 결합한 데이터입니다. 그러나 지역 간의 문화적 다양성과 문화적 차이의 미묘함은 이러한 결합된 평가 데이터 구축에 상당한 어려움을 야기합니다. 이러한 문제를 해결하기 위해, 우리는 권위 있는 문화적 지식 설명을 큐레이션하고, LLM 기반 자동 질의 생성 및 철저한 수동 검증을 통합하는 새로운 프레임워크를 제안합니다. 그 결과, 4.8K개의 수동으로 분해된 세분화된 문화적 설명과 이에 상응하는 48K개의 수동으로 검증된 안전 및 지식 관련 질의를 포함하는 AdaCultureSafe 데이터셋을 확보했습니다. 구축된 데이터셋을 기반으로, 우리는 세 가지 주요 LLM 패밀리에 대한 문화적 안전성과 지식 숙련도를 평가했으며, 이를 통해 중요한 사실을 발견했습니다. 즉, LLM의 문화적 안전성과 지식 숙련도 간에는 유의미한 상관관계가 존재하지 않습니다. 우리는 LLM 내의 유틸리티 관련 뉴런 활성화를 분석하여 상관관계 부재의 잠재적 원인을 조사했습니다. 이는 사전 훈련과 사후 정렬의 목표 차이 때문으로 파악되었습니다. 마지막으로, 우리는 지식 기반 방법을 제시하여 LLM 응답 생성 과정에 지식을 통합함으로써 문화적 안전성을 크게 향상시킵니다.
With the widespread adoption of Large Language Models (LLMs), respecting indigenous cultures becomes essential for models' culturally safety and responsible global applications. Existing studies separately consider cultural safety and cultural knowledge and neglect that the former should be grounded by the latter. This severely prevents LLMs from yielding culture-specific respectful responses. Consequently, adaptive cultural safety remains a formidable task. In this work, we propose to jointly model cultural safety and knowledge. First and foremost, cultural-safety and knowledge-paired data serve as the key prerequisite to conduct this research. However, the cultural diversity across regions and the subtlety of cultural differences pose significant challenges to the creation of such paired evaluation data. To address this issue, we propose a novel framework that integrates authoritative cultural knowledge descriptions curation, LLM-automated query generation, and heavy manual verification. Accordingly, we obtain a dataset named AdaCultureSafe containing 4.8K manually decomposed fine-grained cultural descriptions and the corresponding 48K manually verified safety- and knowledge-oriented queries. Upon the constructed dataset, we evaluate three families of popular LLMs on their cultural safety and knowledge proficiency, via which we make a critical discovery: no significant correlation exists between their cultural safety and knowledge proficiency. We then delve into the utility-related neuron activations within LLMs to investigate the potential cause of the absence of correlation, which can be attributed to the difference of the objectives of pre-training and post-alignment. We finally present a knowledge-grounded method, which significantly enhances cultural safety by enforcing the integration of knowledge into the LLM response generation process.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.