Splash: A Local Engine Built Around the Model (inco.ai)

🤖 AI Summary
In a significant advancement for local AI inference, Inco has launched Splash, an open-source inference engine optimized for Apple silicon. Unlike traditional inference engines that prioritize flexibility over efficiency, Splash is specifically designed around targeted models, enhancing both speed and performance. Currently supporting the Qwen3.8-27B model, Splash boasts a decode speed that is up to 2× faster than competing engines, maintaining this advantage across varying context lengths up to 32K tokens. For users with M3 Macs or newer, it requires a minimum of 36 GB of unified memory, with recommendations for optimal performance at 48 GB. The technical innovations in Splash include specialized kernels, optimized memory management, and a unique batching system that improves cache reuse and response times. By utilizing a hardware-aware memory plan and construction around specific model needs, Splash can efficiently handle multiple concurrent requests—demonstrating a 3.9× throughput boost over the closest alternative when processing multiple simultaneous queries. As a tool that transforms how local inference can be executed, Splash not only elevates individual developer capabilities but also underscores the trend towards highly specialized AI tools tailored to specific use cases. The open-source release invites further collaboration and the expansion of supported models, promising exciting developments ahead for the AI/ML community.
Loading comments...
loading comments...