Unverified benchmarks for the M5 Ultra show 50 tokens per second throughput and 1800 tokens per second prefill for Qwen 2.5 27B (Q4 quantization) at 8k context. The performance was measured without multi-token prediction enabled.
HOW THIS AFFECTS YOU
●
builderThis hardware may offer high-speed local inference for mid-sized models.
●
researcherThese throughput numbers provide a new baseline for evaluating quantized model performance on next-gen silicon.