2608.01651v1 Aug 03, 2026 cs.DC

Bole: 하이브리드 어텐션 언어 모델을 위한 효율적인 트리 추론

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models

Yongchao Liu
Yongchao Liu
Citations: 6
h-index: 1
Yi Su
Yi Su
Soochow University
Citations: 312
h-index: 7
Li Wang
Li Wang
Citations: 0
h-index: 0
Jie Zhang
Jie Zhang
Citations: 0
h-index: 0
Xiabao Wu
Xiabao Wu
Citations: 5
h-index: 1
Chiran You
Chiran You
Citations: 0
h-index: 0
Zhanlong Qiu
Zhanlong Qiu
Citations: 6
h-index: 1
Juelu Zhang
Juelu Zhang
Citations: 0
h-index: 0
Jiajun Zheng
Jiajun Zheng
Citations: 7
h-index: 2
Fangxin Liu
Fangxin Liu
Citations: 0
h-index: 0
Chen Tian
Chen Tian
Citations: 47
h-index: 4
Chengying Huan
Chengying Huan
Citations: 28
h-index: 3

하이브리드 어텐션 대규모 언어 모델은 전체 어텐션과 순환 선형 어텐션을 결합하여 긴 컨텍스트 추론 비용을 줄이지만, 여전히 자기 회귀 디코딩 과정에서 메모리 병목 현상이 발생합니다. 트리 추론은 매력적인 성능 향상 방법이지만, 기존의 트리 추론 시스템은 전체 어텐션 모델의 키-값 캐시에 최적화되어 설계되었습니다. 하이브리드 모델에서는 이러한 시스템이 순환 레이어를 한 분기씩 탐색하며 모든 제안 노드에 대한 전체 상태를 생성하므로, 검증 지연 시간과 임시 메모리 사용량이 트리 크기와 배치 크기에 따라 비효율적으로 증가합니다. 본 논문에서는 하이브리드 어텐션 LLM을 위한 효율적인 트리 추론을 가능하게 하는 커널-런타임 공동 설계인 Bole을 제시합니다. Bole은 선형 어텐션의 순환 구조를 트리 형태의 닫힌 형태로 변환하고, 이를 자원 효율적인 GPU 커널로 구현하여 모든 제안 노드를 병렬로 검증하고 선형 어텐션 트리 검증 속도를 3.4배에서 7.7배까지 향상시킵니다. 또한, Bole은 추론 상태 업데이트를 토큰 레벨의 요소로 무손실 인코딩하고, 샘플링 후 선택된 상태만 재구성하여 임시 상태 메모리 사용량을 82배에서 99배까지 줄이고 KV 캐시에 필요한 GPU 자원을 확보합니다. SGLang이라는 널리 사용되는 LLM 서비스 엔진에 Bole을 통합함으로써 효율적인 상태 관리와 전체 하이브리드 순방향 프로세스에 맞춰 조정된 배치 단위 검증 예산을 결합했습니다. 네 가지 모델, 두 개의 GPU 플랫폼 및 다양한 데이터 세트를 사용하여 Bole은 자기 회귀 디코딩 대비 최대 4.72배 높은 오프라인 디코딩 처리량을 제공하며, 가장 강력한 트리 추론 기반 시스템보다 최대 2.03배 더 빠른 속도를 보입니다. 온라인 에이전트 작업 환경에서는 Bole이 가장 강력한 트리 추론 기반 시스템에 비해 TTFT(Time To First Token)를 최대 67.6% 줄이고 TPOT(Tokens Per Second)를 최대 49.9% 향상시킵니다.

Original Abstract

Hybrid-attention large language models combine full attention with recurrent linear attention to reduce long-context inference costs, yet their autoregressive decoding remains memory-bound. Tree speculative decoding offers an attractive acceleration path, but existing tree-speculation systems are designed around the key--value caches of full-attention models. On hybrid models, they traverse recurrent layers branch by branch and materialize a full state for every proposal node, causing verification latency and transient memory to scale poorly with tree and batch sizes. We present Bole, a kernel--runtime co-design that enables efficient tree speculation for hybrid-attention LLMs. Bole transforms the linear-attention recurrence into a tree-structured closed form and realizes it with a resource-efficient GPU kernel, verifying all proposal nodes in parallel and accelerating linear-attention tree verification by 3.4--7.7$\times$. It losslessly encodes speculative state updates as token-level factors and reconstructs only the state selected after sampling, reducing transient state memory by 82--99$\times$ and freeing GPU capacity for KV caches. Its integration into SGLang, a widely deployed production LLM serving engine, couples efficient state management with a batch-wide verification budget calibrated to the complete hybrid forward. Across four models, two GPU platforms, and diverse datasets, Bole delivers up to $4.72\times$ the offline decode throughput of autoregressive decoding and up to $2.03\times$ that of the strongest tree-speculative baseline. Under online agent workloads, it reduces TTFT and TPOT by up to $67.6%$ and $49.9%$, respectively, over the strongest tree-speculative baseline.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!