[r/LocalLLaMA]score: 0.15
I gave a 21M model a 6.4B-parameter lookup table. It matches a 114M dense model and runs with the table on an SSD (RX 9070)
October 6, 2026
A 21M parameter model utilizing a 6.4B parameter lookup table achieves performance parity with a 114M dense model on 500M Wikipedia tokens. The implementation uses custom Triton kernels to memory-map a 4-bit table from an NVMe SSD, maintaining 140 tokens per second while consuming only 0.4 GB of VRAM.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy