2605.29524v1 May 28, 2026 cs.CR

KBF: 언어 모델 및 블랙박스 API 감사를 위한 지식 경계를 지문으로 활용

KBF: Knowledge Boundary as Fingerprint for Language Model and Black-Box API Auditing

Mingxun Zhou
Mingxun Zhou
Carnegie Mellon University
Citations: 383
h-index: 10
Yijia Fang
Yijia Fang
Citations: 6
h-index: 2
Yiqing Feng
Yiqing Feng
Citations: 10
h-index: 2
Bingyu Li
Bingyu Li
Citations: 6
h-index: 2

리레이 및 리셀러 API가 대규모 언어 모델(LLM)에 대한 접근을 점점 더 중개하고 있지만, 사용자는 특정 엔드포인트가 실제로 광고된 모델을 제공하는지 직접 확인할 수 있는 방법이 없습니다. 본 논문에서는 지식 경계 근처의 안정적인 수치적 재현율을 사용하여 모델 API를 식별하는 저비용 블랙박스 감사 프로토콜인 KBF를 소개합니다. 16개의 실제 LLM 엔드포인트에서 KBF는 경제적으로 중요한 모든 155개의 대체 API를 식별했으며, 동일한 모델에 대한 검증은 모두 통과했습니다. KBF는 배포 변경에도 안정적이며, 전체 트래픽의 5-10%만 대체되는 경우에도 고분산 혼합 라우팅 공격을 탐지할 수 있습니다. 또한, 6개의 플랫폼으로 구성된 숨겨진 API 감사에서 27개의 플랫폼 모델 셀 중 7개가 기준 엔드포인트와 통계적으로 일치하지 않는다는 사실을 발견했으며, 이러한 불일치는 주로 프리미엄 Claude 엔드포인트에서 집중적으로 나타났습니다.

Original Abstract

Relay and reseller APIs increasingly intermediate access to large language models (LLMs), but users have no direct way to verify that a claimed endpoint is actually serving the advertised model. We introduce KBF, a low-cost black-box auditing protocol that fingerprints model APIs using stable numerical recall near the knowledge boundary. Across 16 production LLM endpoints, KBF flags all 155 economically relevant substitutions without rejecting any same-model controls, remains stable under deployment variation, detects high-separation mixed-routing attacks when only 5-10% of traffic is substituted, and finds that 7 of 27 platform model cells in a six-platform shadow API audit are statistically inconsistent with their reference endpoints, with inconsistencies concentrated on premium Claude endpoints.

2 Citations
0 Influential
5 Altmetric
27.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!