소규모 언어 모델의 환각 현상에 대한 기하학적 분석
A Geometric Analysis of Small-sized Language Model Hallucinations
언어 모델, 특히 다단계 또는 에이전트 기반 환경에서, 사실과 다르지만 유창한 응답을 생성하는 '환각' 현상은 언어 모델의 신뢰성을 저해하는 주요 문제입니다. 본 연구에서는 기하학적 관점에서 소규모 언어 모델의 환각 현상을 분석합니다. 모델이 동일한 프롬프트에 대해 여러 응답을 생성할 때, 진실에 기반한 응답들은 임베딩 공간에서 더 촘촘하게 클러스터링된다는 가설을 바탕으로, 본 연구는 이 가설을 입증하고, 이 기하학적 통찰력을 활용하여 일관된 수준의 분리 가능성을 달성할 수 있음을 보여줍니다. 이러한 결과는 30~50개의 주석만으로도 대규모 응답 집합을 분류하는 효율적인 라벨 전파 방법을 제시하며, F1 점수가 90%를 초과하는 성능을 달성합니다. 본 연구는 임베딩 공간에서 환각 현상을 기하학적 관점에서 조명함으로써 기존의 지식 중심적이고 단일 응답 평가 방식에 대한 보완적인 연구 결과를 제공하며, 추가적인 연구를 위한 기반을 마련합니다.
Hallucinations -- fluent but factually incorrect responses -- pose a major challenge to the reliability of language models, especially in multi-step or agentic settings. This work investigates hallucinations in small-sized LLMs through a geometric perspective, starting from the hypothesis that when models generate multiple responses to the same prompt, genuine ones exhibit tighter clustering in the embedding space, we prove this hypothesis and, leveraging this geometrical insight, we also show that it is possible to achieve a consistent level of separability. This latter result is used to introduce a label-efficient propagation method that classifies large collections of responses from just 30-50 annotations, achieving F1 scores above 90%. Our findings, framing hallucinations from a geometric perspective in the embedding space, complement traditional knowledge-centric and single-response evaluation paradigms, paving the way for further research.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.