Ling-3.0-tiny 8B MoE Achieves 105 Tokens/s on DGX Spark
August 10, 2026
Ling-3.0-tiny is an 8B parameter Mixture-of-Experts model with 1.3B active parameters, delivering 86-90 tokens/s on M4 Pro MacBooks and utilizing 8.34 GiB peak memory at 8K context.
HOW THIS AFFECTS YOU
●
builderYou can deploy highly efficient, high-throughput models on consumer-grade hardware.
●
researcherYou can study the performance characteristics of small-scale MoE architectures in competitive benchmarks.