2608.08485v1 Aug 09, 2026 cs.AI

HoloAegis: 동결된 표현과 위상 추론: 제로샷 LLM 안전 장벽을 위한 최소 매개변수 기반 안전 영역

HoloAegis: Frozen Representation, Topological Inference: Minimally Parametric Safety Manifolds for Zero-Shot LLM Guardrails

Lik-Hang Lee
Lik-Hang Lee
Citations: 1,029
h-index: 10
Tak Ho Alex Li
Tak Ho Alex Li
Citations: 0
h-index: 0
Kaijie Liu
Kaijie Liu
Citations: 0
h-index: 0
Kin Chung Ho
Kin Chung Ho
Citations: 0
h-index: 0
Ping Shum
Ping Shum
Citations: 0
h-index: 0
Michael K. Ng
Michael K. Ng
Citations: 0
h-index: 0

현재 LLM 안전 장벽은 근본적인 어려움에 직면합니다. 미세 조정은 사전 학습된 표현을 왜곡시키고, 생성적 판단기는 엄청난 추론 비용을 발생시킵니다. 우리는 기존의 패러다임을 뒤엎으며 다음과 같은 질문을 던집니다: 동결된 의미 표현에 대한 순수한 기하학적 추론만으로 안전성을 달성할 수 있을까요? 우리는 HoloAegis를 제안합니다. HoloAegis는 표현과 추론을 분리하는 최소 매개변수 기반의 위상 추론 프레임워크입니다. 우리의 접근 방식을 최소 매개변수 기반이라고 부르는 이유는 자유 매개변수가 앵커 개수 K와 온도 파라미터 tau뿐이며, 이는 구축 후 고정되어 있으며 그래디언트 기반 학습을 필요로 하지 않기 때문입니다. 미세 조정되지 않은 인코더는 텍스트를 단위 구에 매핑하고, 이후 모든 결정은 순전히 기하학적입니다. 우리는 안전성 평가를 사전 계산된 시스템 위상 앵커 은행에 대한 Gibbs-Boltzmann 자유 에너지 계산으로 공식화하며, 점진적인 다중 턴 의미 변화를 감지하기 위해 이중 시간 규모 지수 이동 평균을 도입합니다. 우리의 핵심 이론적 통찰력은 '위상 경계 안정성 가설'입니다. 우리는 이론적 근거와 강력한 경험적 증거를 제시하여 희소 앵커 중심점이 전체 벡터 공간 방법보다 고주파 어휘 변화에 대한 결정 경계를 훨씬 더 잘 안정화시킨다는 것을 보여줍니다. HoloAegis는 8개의 벤치마크에서 평가되었으며, 최고 수준의 정확도(AuthenHallu에서 1.0000 AUC, HarmBench에서 0.9802)를 달성했으며, 서브 밀리초 지연 시간, 초기 데이터 불필요, 그리고 교차 언어 전이 기능(중국어 CHIFRAUD에서 0.9758 AUC)을 제공합니다.

Original Abstract

Current LLM safety guardrails face a fundamental tension: fine-tuning distorts pre-trained representations while generative judges incur prohibitive inference costs. We challenge the prevailing paradigm by asking: can safety be achieved through pure geometric reasoning over frozen semantic representations? We present HoloAegis, a minimally parametric topological inference framework that decouples representation from reasoning. We term our approach minimally parametric because the only free parameters are the anchor count K and the temperature tau, both fixed after construction and requiring no gradient-based training. An un-fine-tuned encoder maps text to a unit sphere, after which all decisions are purely geometric. We formalize safety evaluation as a Gibbs-Boltzmann Free Energy computation over a pre-computed System Topology Anchor Bank, and we introduce Dual Time-Scale Exponential Moving Averages to detect progressive multi-turn semantic drift. Our key theoretical insight is a Topological Boundary Stability Conjecture: we provide theoretical motivation and strong empirical evidence that sparse anchor centroids stabilize the decision boundary against high-frequency lexical perturbations far better than full vector space methods. Evaluated across 8 benchmarks, HoloAegis achieves state-of-the-art accuracy (1.0000 AUC on AuthenHallu, 0.9802 on HarmBench) with sub-millisecond latency, zero cold-start data, and cross-lingual transfer (0.9758 AUC on Chinese CHIFRAUD).

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!