Qwen3.8-27B at ~200 tok/s peak on an Apple M5 Max (twitter.com)

🤖 AI Summary
Lithos.ai has announced the open-sourcing of lithos-metal, a cutting-edge technology that leverages megakernels and DSpark speculative decoding to achieve impressive performance. The Qwen3.8-27B model can now process at a peak rate of over 200 tokens per second on an Apple M5 Max. This development marks a significant advancement in the usability of AI tools, allowing users to run ultra-fast inference directly on their laptops with a simple command, thus democratizing access to powerful AI capabilities. For the AI/ML community, this innovation is noteworthy as it bridges the gap between powerful, complex models and accessible, real-time processing on consumer-grade hardware. The ability to execute intensive AI tasks locally not only enhances performance but also reduces reliance on cloud computing resources, which can introduce latency and cost concerns. By making high-performance AI tools more accessible, lithos-metal could foster greater experimentation and application of AI technologies in various fields. More technical details and insights can be found in their blog post and on their GitHub repository.
Loading comments...
loading comments...