DeepSeek V4 Flash, up to 32 tok/s on AMD Ryzen AI MAX+ 395 (www.lucebox.com)

🤖 AI Summary
DeepSeek has announced the launch of V4 Flash, a powerful 284-billion-parameter mixture-of-experts model that achieves an impressive decoding speed of up to 32 tokens per second (tok/s) on the AMD Ryzen AI MAX+ 395. Leveraging the integrated capabilities of the CPU and the Radeon 8060S GPU, this model runs entirely local without the need for discrete GPUs or remote inference services, showcasing significant advancements in efficiency for AI operations. Moreover, through its innovative ROCmFPX format, which allows efficient memory utilization, the model is capable of executing complex tasks that require dynamic routing of expert networks. This development marks a significant milestone for the AI/ML community by pushing the boundaries of locally running large models, providing a competitive edge in performance metrics compared to previous benchmarks such as HipFire and DwarfStar. The incorporation of advanced techniques like indexed sparse prefill—achieving around 250 tok/s—and a tailored HIP decode path enhances the model's speed and responsiveness. The experiments demonstrate how optimized hardware-software integration can significantly improve AI processing speeds while maintaining quality, paving the way for faster and more powerful local inference solutions.
Loading comments...
loading comments...