I ran a 397B-parameter model on 20 GPUs with 16 GB each (diljitpr.net)

🤖 AI Summary
A groundbreaking experiment has successfully run the massive 397 billion-parameter Qwen3.5 model on a decentralized network of 20 GPUs, each with only 16 GB of memory, using a novel open-source project called Sangama. By employing pipeline parallelism, the model was split into stages distributed across the machines, demonstrating that ordinary home computers can collectively handle large-scale AI models without the need for high-bandwidth data center setups. Initial tests revealed significant network latency, limiting throughput to just 0.6 tokens per second, but innovative changes improved performance to between 5 and 5.5 tokens per second by optimizing data transfer and reducing computation times. This achievement is significant for the AI/ML community as it opens the door for larger models to be processed on smaller, more accessible hardware, rather than relying solely on expensive machinery. Moreover, the approach demonstrated the importance of geographic proximity in network performance and highlighted potential optimizations, such as minimizing data transfer costs and supporting multiple concurrent requests. The experiment not only enhances our understanding of model distribution but also invites further exploration into cost-effective, collaborative AI advancements. The project's code is available on GitHub, encouraging community participation in this innovative frontier of decentralized AI processing.
Loading comments...
loading comments...