Daedalus-150M Hybrid Architecture Optimized for CPU Inference
August 19, 2026
Daedalus-150M uses a hybrid convolution-attention architecture designed specifically for 4-bit CPU inference. By using short convolutions in 12 of its 18 blocks, it limits memory usage to two timesteps, outperforming GPT-2 124M and MobileLLM-125M despite being trained on fewer tokens.
HOW THIS AFFECTS YOU
●
builderYou can deploy highly efficient 150M parameter models on standard CPUs with limited memory overhead.
●
founderThis architecture enables high-performance AI features on low-cost, edge-computing hardware.