Petals
petals.dev 1 month ago 35
Petals enables users to run large language models like Llama 3.1 and Mixtral on consumer-grade hardware by distributing model layers across a peer-to-peer network similar to BitTorrent. The system achieves inference speeds of up to 6 tokens per second for Llama 2 (70B) and supports fine-tuning and custom model paths through PyTorch. This approach makes running billion-parameter models accessible to individuals without enterprise-grade infrastructure.