2605.25645v1 May 25, 2026 cs.DC

구글 클라우드 TPU에서 Gemma 4 31B 모델을 미세 조정하고 서비스하는 방법: GPU 기반 시스템과의 기술적 비교

Fine-Tuning and Serving Gemma 4 31B on Google Cloud TPU: A Technical Comparison with GPU Baselines

Pulkit Agrawal
Pulkit Agrawal
Citations: 356
h-index: 7
Jatin Kishnani
Jatin Kishnani
Citations: 0
h-index: 0
Mayank Goel
Mayank Goel
Citations: 8
h-index: 2
A. Singh
A. Singh
Citations: 8
h-index: 2
Sairanjan Mishra
Sairanjan Mishra
Citations: 0
h-index: 0

본 논문에서는 구글의 Gemma 4 31B 모델을 TPU 하드웨어에 적용하여 미세 조정 및 서비스를 수행하는 첫 번째 완전한 사례 연구를 제시하며, 대규모 언어 모델 적응을 위한 TPU와 GPU 플랫폼 간의 실증적 비교 결과를 제공합니다. PyTorch, Hugging Face TRL 및 FSDP 기반으로 구축된 GPU용 학습 레시피를 JAX + Tunix/Qwix 스택으로 이식하기 위해 필요한 모든 코드 수준 변경 사항을 기록했습니다. 이러한 변경 사항에는 메쉬 구성, LoRA 모듈 명명 규칙, 샤딩 주석 수정, 그래디언트 체크포인팅, 데이터 파이프라인 재구성 및 Orbax에서 safetensors로의 사용자 정의 체크포인트 병합 절차가 포함됩니다. 추론 과정에서는 Gemma 4를 v6e-8에서 서비스하기 위해 필요한 vLLM-TPU Docker 설정을 상세히 설명하고, 결과적인 지연 시간과 처리량 프로파일을 분석합니다. 동일한 하이퍼파라미터 환경에서 2개의 H100 GPU를 사용한 기준 시스템과 비교했을 때, TPU 학습은 1.61배 더 빠르게 완료되고 비용은 2.12배 낮습니다. 추론 처리량은 플랫폼 간에 3% 이내의 차이를 보이지만, TPU는 첫 번째 토큰 생성까지 걸리는 시간이 2배 짧습니다 (235ms vs. 475ms). 종합적으로, TPU 구성은 대표적인 학습 및 서비스 워크로드에서 1.82배 더 저렴합니다. 본 연구는 공개 도구 생태계의 중요한 격차를 해소하고, 실무자들이 Gemma 4 모델을 TPU 인프라에 배포하기 위한 재현 가능하고 프로덕션 환경에 적합한 레시피를 제공합니다.

Original Abstract

We present the first end-to-end demonstration of fine-tuning and serving Google's Gemma 4 31B model on TPU hardware, providing an empirical comparison of TPU and GPU platforms for large language model adaptation. Using LoRA on a Google TPU v5p-8 for training and TPU v6e-8 (Trillium) for inference, we document the full set of code-level adaptations required to port a GPU-native training recipe, built on PyTorch, HuggingFace TRL, and FSDP, to the JAX + Tunix/Qwix stack. These adaptations span mesh configuration, LoRA module naming conventions, sharding annotation corrections, gradient checkpointing, data pipeline restructuring, and a custom Orbax-to-safetensors checkpoint merging procedure. For inference, we detail the vLLM-TPU Docker setup necessary to serve Gemma 4 on v6e-8 and characterize the resulting latency and throughput profile. Compared with a 2xH100 GPU baseline under identical hyperparameters, TPU training completes 1.61x faster at 2.12x lower cost. Inference throughput is within 3% across platforms, while TPU achieves 2x lower time-to-first-token (235 ms vs. 475 ms). Together, the TPU configuration is 1.82x cheaper for a representative train-plus-service workload. Our work removes a critical gap in the open tooling ecosystem and provides practitioners with a reproducible, production-ready recipe for Gemma 4 deployment on TPU infrastructure.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!