Someone runs K3 on 80x 5090s, for 20 tok/s (twitter.com)

🤖 AI Summary
Kimi has successfully deployed the K3 model, boasting 2.8 trillion parameters, on a fleet of 80 RTX 5090 GPUs, achieving an impressive throughput of 20 tokens per second on the first day of operation without any tuning. This deployment marks a significant milestone as it utilizes standard GDDR7 graphics cards rather than high-bandwidth memory (HBM), which is typically scarce in the AI field. Previously, the team increased the performance of the GLM-5.2 model from 30 to 110 tokens per second with the same GPU setup, indicating that the K3's performance is likely to improve as optimizations are implemented. This achievement underscores a pivotal moment for open-weight AI models, providing unprecedented access to advanced technology for labs, startups, and educational institutions. With K3 being heralded as the most powerful open model currently available, the implications are vast; researchers and developers are empowered to experiment, fine-tune, and integrate cutting-edge AI capabilities into their projects using widely available resources. The accessibility of high-performance AI on consumer-grade GPUs could democratize innovation in the AI/ML community, opening doors for a broader range of applications and studies.
Loading comments...
loading comments...