2607.05708v1 Jul 07, 2026 cs.AI

Akashic: MemAttention을 이용한 저오버헤드 LLM 추론 서비스

Akashic: A Low-Overhead LLM Inference Service with MemAttention

Junhao Hu
Junhao Hu
Citations: 196
h-index: 5
Chentao Wu
Chentao Wu
Citations: 155
h-index: 6
Zhaokai Luo
Zhaokai Luo
Citations: 0
h-index: 0
Huayi Jin
Huayi Jin
Citations: 1
h-index: 1
Zhiyong Wang
Zhiyong Wang
Citations: 6
h-index: 1
Ruozhou He
Ruozhou He
Citations: 5
h-index: 1
Chenchen Hong
Chenchen Hong
Citations: 0
h-index: 0
Yunfei Gu
Yunfei Gu
Citations: 27
h-index: 3
Yang Liu
Yang Liu
Citations: 0
h-index: 0
Yifei Liu
Yifei Liu
Citations: 6
h-index: 2

최근의 LLM 기반 에이전트 시스템은 다중 대화, 도구 호출 및 세션 간 워크플로우를 통해 지속적으로 문맥 정보를 축적합니다. 모든 요청에 대해 전체 기록을 재생하는 것은 빠르게 비실용적이 되며, 긴 문맥은 프리필 비용을 증가시키고, 컨텍스트 제한을 초과할 수 있으며, 종종 관련 없는 내용 속에 중요한 정보가 묻혀 서비스 효율성과 출력 품질을 저하시킵니다. 본 논문에서는 MemAttention을 기반으로 설계된 저오버헤드 메모리 시스템인 Akashic을 제안합니다. Akashic은 문맥 정보를 경계화된 청크(chunk)로 구성하고, 청크 간의 의미적 관계를 모델링하여 전체 기록을 반복적으로 다시 쓰지 않고도 청크 간의 연결성을 유지합니다. 또한, Akashic은 하드웨어-소프트웨어 공동 설계 메모리 배치 기술을 적용하여 함께 검색될 가능성이 높은 청크들을 함께 위치시켜 검색 단편화 및 I/O 오버헤드를 줄입니다. 네 가지 대표적인 워크로드와 세 가지 모델 크기에서 실험한 결과, Akashic은 기존의 강력한 메모리 기반 시스템에 비해 작업 정확도를 최대 10.2 포인트, 처리량을 최대 1.21배, 지속 가능한 요청률을 최대 1.88배 향상시켰습니다.

Original Abstract

Recent LLM-based agent systems continuously accumulate context across multi-turn interactions, tool invocations, and cross-session workflows. Replaying the full history for every request quickly becomes impractical: long contexts increase prefill cost, may exceed context limits, and often bury task-relevant evidence in irrelevant content, degrading both serving efficiency and output quality. We propose Akashic, a low-overhead memory system built around MemAttention, which organizes context into bounded chunks and models semantic relationships across chunks, preserving cross-chunk evidence without repeatedly rewriting the full history. Akashic further applies hardware-software co-designed memory placement to co-locate likely co-retrieved chunks, reducing retrieval fragmentation and I/O overhead. Across four representative workloads and three model sizes, Akashic improves task accuracy by up to 10.2 points, throughput by up to 1.21x, and sustainable request rate by up to 1.88x over strong prior memory baselines.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!