LFM 2.6B achieves high throughput on consumer hardware
August 9, 2026
The LFM 2.6B model enables high-speed inference on single RTX 3090 GPUs, reaching 260 tokens per second. It is optimized for low-latency tasks like summarization and command autocomplete with a 128k context window.
HOW THIS AFFECTS YOU
●
builderYou can deploy this model for high-throughput, low-latency edge tasks or mobile-first applications.