vllm: FP8 KV cache for Triton DiffKV and MiMo-V2.6-Flash
October 6, 2026
vLLM integration adds support for Triton DiffKV and MiMo-V2.6-Flash models, enabling FP8 quantization for the KV cache. This update allows practitioners to reduce memory overhead during long-context inference by utilizing lower-precision key-value storage.