AngelSpec Unifies MTP and Block-Parallel Speculative Decoding
July 29, 2026
AngelSpec co-specializes speculative decoding architectures and data to handle heterogeneous workloads. It uses multi-token prediction for high-entropy chat and block-parallel diffusion for predictable code and math sequences to optimize inference speeds.
HOW THIS AFFECTS YOU
●
builderYou can achieve higher inference throughput by matching drafting structures to specific data distributions.
●
researcherYou can leverage specialized training for different drafting mechanisms.