Bonsai: A 27B reasoning model on a 16 GB M2 Mac, with Ferrox (medium.com)

🤖 AI Summary
PrismML has introduced Bonsai, a groundbreaking 27 billion parameter reasoning model capable of running on a 16 GB M2 Mac thanks to the new Ferrox inference engine. Ferrox, developed in Rust, enables compact model execution through a unique ternary quantization scheme that packs weights efficiently at 1.75 bits per weight without sacrificing significant performance. Bonsai demonstrates remarkable efficiency, retaining 98.2% of FP16-level intelligence while consuming only 5.95 GB of disk space, allowing for high-caliber reasoning on consumer-grade hardware. The significance of Bonsai lies in its advanced quantization method, leveraging a Walsh-Hadamard transformation to enhance the model's performance despite the reduced precision. This ternary structure allows Bonsai to tackle complex reasoning tasks while maintaining a minimal footprint, making powerful AI more accessible. With an impressive average score of 84.78 across various benchmarks, Bonsai's performance rivals much larger models, showing that substantial efficiency gains can be achieved without compromising output quality. Additionally, users can easily deploy Ferrox to run Bonsai locally, fostering greater accessibility for developers interested in cutting-edge AI capabilities.
Loading comments...
loading comments...