Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs (github.com)

🤖 AI Summary
The Kimi K3 model, a 2.8-trillion-parameter AI, can now be run on a MacBook Pro at a rate of 1 token per second, thanks to the Deltafin project. This work is built on a fork of gavamedia/deltafin and utilizes advanced storage techniques via the ARGODRIVE system to streamline the model's performance directly from four SSDs, without compromising the integrity of its massive parameter set. Deltafin emphasizes a no-pruning approach, ensuring that the model operates at its full potential, with each token generated being validated by the model itself, resulting in consistently high-quality outputs. This development is significant for the AI/ML community as it demonstrates that sophisticated models like Kimi K3 can be executed on consumer-grade hardware, drastically reducing the cost and accessibility barriers for developers and researchers. The project aims to extract maximum efficiency from existing infrastructure, showcasing the feasibility of high-performance AI without requiring expensive setups. It serves as a benchmark for future experiments in running large models locally and could inform advancements in scalable AI applications, reinforcing the notion that cutting-edge AI can be more accessible to a broader audience while maintaining rigorous quality standards.
Loading comments...
loading comments...