2606.23590v1 Jun 22, 2026 cs.AI

부적절한 질문의 위상 구조: LLM에서의 탐지 및 제어를 위한 지속적인 호몰로지

The Topology of Ill-Posed Questions: Persistent Homology for Detection and Steering in LLMs

Sizhe Tang
Sizhe Tang
Citations: 32
h-index: 4
Tian Lan
Tian Lan
Citations: 28
h-index: 3
Mahdi Imani
Mahdi Imani
Citations: 107
h-index: 6
Guangyu Jiang
Guangyu Jiang
Citations: 18
h-index: 2

모호하거나, 불완전하게 정의되었거나, 모순되는 질문과 같이 부적절한 질문은 유효한 답변을 전혀 제공하지 않거나 여러 개의 가능한 답변을 제시할 수 있으며, 이는 대규모 언어 모델(LLM)에게 어려움을 야기합니다. 기존의 접근 방식은 주로 모델의 출력 결과를 분석하며, 종종 특정 하위 클래스에 초점을 맞춥니다. 본 연구에서는 다양한 유형의 부적절성을 LLM의 내부 상태의 통합된 위상 구조 내에서 표현할 수 있는지, 그리고 이러한 구조가 응답 동작을 제어하는 데 사용될 수 있는지 조사합니다. 프롬프트 토큰의 컨텍스트 임베딩을 각 트랜스포머 레이어에서 점군으로 모델링하고, 유한 차원의 지속적인 호몰로지를 사용하여 이 점군의 기하학적 특성을 분석합니다. 각 레이어는 평균 유한 수명, 정규화된 수명 엔트로피 및 최대 수명 농도를 나타내는 세 가지 주요 지표로 요약됩니다. 이러한 지표들을 레이어별로 연결하면 질문의 위상 표현을 얻을 수 있습니다. 또한, 우리는 위상 정보를 활용하여 유사한 예제를 검색하고, 질문에 특정한 활성화 조작을 수행함으로써 출처를 인식하는 명확화 또는 회피를 유도하는 '위상 기반 활성화 제어' 방법을 도입합니다. 세 가지 공개 LLM 모델에서, 위상 특징은 부적절성 분류 작업에서 프롬프트 기반 및 풀링된 숨겨진 상태 방법보다 일관되게 우수한 성능을 보였습니다. AmbigQA 데이터셋에서 평균 정확도가 67.4%에서 78.9%로 향상되었고, SituatedQA 데이터셋에서는 79.9%에서 88.5%, CLAMBER 9-way 분류 작업에서는 57.6%에서 69.6%로 개선되었습니다. 위상 기반 제어는 평균 허용 가능한 응답률을 61.4%에서 70.6%로, 그리고 근거 있는 허용 가능한 응답률을 11.9%에서 16.4%로 증가시켰습니다. 이러한 결과는 지속적인 호몰로지가 부적절성을 해석 가능하게 표현하고, 대상 응답 제어를 위한 효과적인 메커니즘을 제공한다는 것을 보여줍니다.

Original Abstract

Ill-posed questions, including ambiguous, underspecified, or contradictory queries, may admit no valid answer or multiple plausible answers, posing a challenge for large language models (LLMs). Existing approaches largely analyze ill-posedness through model outputs and often focus on specific subclasses. We investigate whether diverse sources of ill-posedness can be represented within a unified topology of LLM internal states and whether this structure can be used to steer response behavior. We model the contextual hidden states of prompt tokens at each transformer layer as a point cloud and characterize its geometry using finite zero-dimensional persistent homology. Each layer is summarized by three compact descriptors: mean finite lifetime, normalized lifetime entropy, and largest-lifetime concentration. Concatenating these descriptors across layers yields a topology representation of the question. We further introduce topology-conditioned activation steering, which retrieves topologically similar examples and constructs query-specific activation interventions that encourage source-aware clarification or abstention. Across three open-weight LLMs, topology features consistently outperform prompt-based and pooled-hidden-state baselines for ill-posedness classification, improving average accuracy from \(67.4\%\) to \(78.9\%\) on AmbigQA, from \(79.9\%\) to \(88.5\%\) on SituatedQA, and from \(57.6\%\) to \(69.6\%\) on CLAMBER 9-way classification. Topology-conditioned steering increases the average total acceptable response rate from \(61.4\%\) to \(70.6\%\) and grounded acceptable responses from \(11.9\%\) to \(16.4\%\). These results show that persistent homology provides both an interpretable representation of ill-posedness and an effective mechanism for targeted response steering.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!