2605.30040v1 May 28, 2026 cs.CR

토큰 인플레이션: 비윤리적인 제공업체가 대규모 언어 모델 사용량에 대해 과다 청구하는 방법

Token Inflation: How Dishonest Providers Can Overcharge for Large Language Model Usage

Jinghuai Zhang
Jinghuai Zhang
Citations: 55
h-index: 3
Shahinul Hoque
Shahinul Hoque
Citations: 8
h-index: 1
Fnu Suya
Fnu Suya
Citations: 392
h-index: 9
Jinyuan Sun
Jinyuan Sun
Citations: 5
h-index: 1

대규모 언어 모델(LLM)의 상용화된 가격 정책에서 토큰당 요금제가 일반적이며, 보고되는 토큰 수의 정확성은 사용자가 지불해야 하는 금액에 직접적인 영향을 미칩니다. 본 연구는 이러한 방식의 과금 시스템이 설계상 감사하기 어렵다는 점을 보여줍니다. 제공업체들은 모델 자체, 토크나이저, 실행 과정을 숨김으로써 지적 재산을 보호하고, 악용 시도를 방지하며, 사용자 개인 정보를 보호합니다. 따라서 감사는 제공업체가 제공하는 증거만을 검토하게 됩니다. 결과적으로, 감사는 제공업체의 보고 내용에 대한 일관성 검토로 귀결됩니다. 우리는 이를 '신뢰의 역설'이라고 부릅니다. 모든 감사 과정은 어떤 형태로든 신뢰를 전제로 해야 하지만, 현재의 시스템은 제공업체가 조작할 가장 큰 동기를 가진 요소에 대한 신뢰를 요구합니다. 본 연구에서는 최근 개발된 세 가지 토큰 감사 프레임워크를 분석하고, 일반적인 상업적 능력을 갖춘 제공업체가 체계적으로 청구되는 토큰 수를 늘릴 수 있음을 보여줍니다. 가장 관대한 설정에서, 숨겨진 추론 과정의 사용량이 평균 146.9%까지 검출되지 않고 과다 청구될 수 있습니다. 현재 최고 수준의 추론 서비스 가격으로 계산하면, 100달러의 정당한 요금이 동일한 질의에 대해 약 1569달러로 부풀려질 수 있습니다. 사용자가 전체 추론 과정을 확인할 수 있더라도, 토큰화 과정의 모호성만으로도 검출 기준 이하에서 평균 50.85%까지 과다 보고가 발생할 수 있습니다. 이러한 결과는 특정 감사 시스템의 문제라기보다는, 증거가 감사 대상 당사자로부터 제공되는 모든 감사 방식에 내재된 문제임을 시사합니다. 정직한 과금 방식을 회복하기 위해서는 신뢰할 수 있는 실행 인증, 추론 과정에 대한 암호화 증명 또는 제3자 재실행과 같이, 제공업체가 통제하지 않는 증거를 기반으로 토큰 수를 검증하는 방법이 필요합니다.

Original Abstract

Per-token billing is now the standard pricing model for commercial large language models (LLMs), so the honesty of reported token counts directly affects what users pay. We show that this kind of billing is hard to audit by design: providers hide the model, the tokenizer, and the execution to protect their IP, mitigate jailbreaks, and preserve user privacy, which means an auditor can only inspect proofs the provider supplies. The audit therefore reduces to a consistency check on the provider's own reports. We call this a trust paradox: every audit must trust some artifact, but current frameworks trust exactly the ones a provider has the strongest reason to manipulate. We study three recent token auditing frameworks and show that a provider with ordinary commercial capabilities can systematically inflate billed token counts. In the most permissive setting, hidden reasoning usage can be inflated by 1,469% on average without detection. At current frontier reasoning prices, that turns a \$100 honest bill into roughly a \$1,569 bill on the same query. Even when the user can see the full reasoning string, tokenization ambiguity alone still allows 50.85% over-reporting below the detection threshold. These results suggest the problem is not in any specific auditor but in any audit whose evidence comes from the audited party. Restoring honest billing will require verification that ties reported token counts to evidence the provider does not control, such as trusted execution attestation, cryptographic proofs of inference, or third-party re-execution.

1 Citations
0 Influential
4.5 Altmetric
23.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!