Petals enables distributed inference for Llama 3.1 405B via P2P network
July 22, 2026
Petals allows users to run large models like Llama 3.1 405B and Mixtral 8x22B by joining a BitTorrent-style network of consumer GPUs. Single-batch inference reaches 6 tokens/sec for Llama 2 70B, offering PyTorch-compatible fine-tuning and custom sampling methods without high-end server hardware.
HOW THIS AFFECTS YOU
●
builderYou can run massive models using consumer hardware by distributing weights across a P2P network.
●
founderYou can bypass high cloud inference costs for large-scale models by leveraging decentralized compute.