Show HN: Slotstream, run Qwen3.8-Flash-Next 4-bit on a low-memory Mac (github.com)

🤖 AI Summary
Slotstream has launched a groundbreaking tool that enables users to run the Qwen3.8-Flash-Next AI model—a resource-intensive, 104 GB model—on lower-memory Macs, starting at just 8.1 GB of RAM by streaming it from an SSD. This is significant for the AI/ML community as it democratizes access to powerful AI tools, allowing more users, especially those with limited hardware, to experiment with advanced models without the need for high-end machines. The Swift binary integrates seamlessly with popular endpoints like Ollama and OpenAI’s SDK, making it user-friendly and accessible. Key technical details include the model's operation parameters: a 48 GB Mac can achieve a warm decode rate of 12 tokens per second, while a cold start takes about 3 seconds. Users must have around 110 GB of free disk space for the model weights, making a 512 GB SSD the practical minimum. The software boasts features such as caching and resizing to optimize performance in real-time, adapting to the user's available system resources. Furthermore, it supports multiple prompt configurations while ensuring efficient memory usage, which can significantly enhance the experience of running large models on machines that would typically struggle to do so.
Loading comments...
loading comments...