[arXiv]score: 0.24
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
October 9, 2026
Mooncake implements a disaggregated architecture for LLM serving by separating prefill and decoding clusters to optimize KVCache management. It utilizes CPU, DRAM, and SSD resources for a distributed cache, achieving up to 525% throughput increases in long-context simulations while maintaining latency SLOs through a prediction-based early rejection policy.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy