2606.19857v1 Jun 18, 2026 cs.CL

대규모 언어 모델은 항상 사람이 읽기 쉬운 언어를 필요로 하지 않는다

Large Language Models Do Not Always Need Readable Language

Hao Peng
Hao Peng
Citations: 28
h-index: 2
Jiayi Zhu
Jiayi Zhu
Citations: 0
h-index: 0
Junxi Wang
Junxi Wang
Citations: 22
h-index: 3
Liang Ke
Liang Ke
Citations: 0
h-index: 0
Chen Zhang
Chen Zhang
Citations: 27
h-index: 3
Linfeng Zhang
Linfeng Zhang
Citations: 492
h-index: 8

대규모 언어 모델(LLM)은 종종 사람에게 읽히기 쉬운 자연어로 프롬프트되고 인터페이스되는데, 심지어 대상 독자가 다른 모델일 때에도 마찬가지입니다. 본 논문에서는 의미 정보가 사람이 읽기 쉽지 않지만 LLM에 의해 복구 가능한 간결하고 표준화되지 않은 텍스트 형태로 인코딩될 수 있는지 조사합니다. 우리는 이러한 모델 중심 텍스트 표현 방식을 '바벨텔레(BabelTele)'이라고 부르며, 여기서는 고정된 프로토콜이 아닌 LLM의 그러한 표현을 생성하고 해석하는 능력을 탐구하기 위한 경험적 도구로 접근합니다. 가독성 진단, 모델 신뢰도 측정, 설문 조사 및 후속 작업 평가를 통해, 바벨텔레가 일반적인 자연어와 크게 달라질 수 있지만 명령 튜닝된 LLM의 경우 핵심 의미를 유지할 수 있음을 확인했습니다. 작업에 독립적인 표현 패러다임으로서 바벨텔레는 높은 정보 밀도를 보여주며, 텍스트 길이가 원래 길이의 27.9%로 축소되더라도 99.5%의 의미 충실도를 유지합니다. 또한, 모델 간 전송, 에이전트 메모리 및 다중 에이전트 통신에서 바벨텔레의 의미적 견고성을 평가했습니다. 결과는 바벨텔레가 컨텍스트 오버헤드를 줄이는 동시에 일반적으로 안정적인 후속 성능을 유지할 수 있지만, 그 효과는 압축기-리더 쌍과 작업 설정에 따라 달라진다는 것을 시사합니다. 이러한 연구 결과는 인간의 가독성, 자연어의 일반성 및 모델 측면에서의 의미 복구 가능성이 부분적으로 분리될 수 있음을 보여주며, 향후 LLM 시스템 탐색을 위한 모델 자체 표현 방식을 개발하는 데 기여할 수 있습니다.

Original Abstract

Large language models (LLMs) are commonly prompted and interfaced with human-readable natural language, even when the intended reader is another model. This paper investigates whether semantic information can be encoded in compact, non-standard textual forms that sacrifice human readability while remaining recoverable by LLMs. We refer to this class of model-centric textual representations as BabelTele, approached here not as a fixed protocol but as an empirical probe into LLMs' capacity to generate and interpret such representations. Through readability diagnostics, model likelihood measures, human questionnaires, and downstream task evaluations, we find that BabelTele can substantially depart from ordinary natural language while preserving core semantics for instruction-tuned LLMs. As a task-agnostic representational paradigm, BabelTele demonstrates high information density, maintaining 99.5% semantic fidelity even when the text volume is condensed to 27.9% of its original length. We further evaluate its semantic robustness in cross-model transfer, agent memory, and multi-agent communication. Results suggest that BabelTele can reduce context overhead while generally maintaining reliable downstream performance, although its effectiveness depends on the compressor-reader pair and task setting. These findings indicate that human readability, natural-language typicality, and model-side semantic recoverability can be partially decoupled, opening a path toward model-native representations in future exploration of LLM systems.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!