A Kimi-k3 GGUF implementation via llama.cpp achieved 0.41 tok/s prompt evaluation and 0.23 tok/s generation on dual RTX 6000 PRO hardware. The setup utilized a 9965WX PRO system with 512 GB DDR5 to handle the model weights.
HOW THIS AFFECTS YOU
●
builderYou can use these benchmarks to estimate hardware requirements for running Kimi-k3 locally.
●
researcherYou can reference these throughput numbers for comparative analysis of large-scale MoE model inference.