Gufo-Qwen3.6-35B-A3B-Q6dense – 3095tok/s prefill; 190 tok/s decode on Strix Halo (github.com)

🤖 AI Summary
Gufo has announced the launch of its vertical local inference engine, specifically designed to optimize and run on the AMD Strix Halo hardware ecosystem, which includes systems equipped with the Ryzen AI MAX+ 395 and Radeon 8060S graphics card. The engine boasts impressive performance metrics, recording a peak of 3,095 tokens per second for prefill and 190 tokens per second for decoding with the robust Qwen3.6 model. This development signifies a crucial step for the AI/ML community by enabling high-performance local inference capabilities tailored for AMD's hardware, thus potentially enhancing access to advanced AI models for developers and researchers. The Gufo platform supports several advanced AI models across different modalities, including text, audio, image, and video. It prioritizes speed and efficiency by focusing on models that can operate within a unified memory space of up to 128 GiB, ensuring high-quality performance even on smaller configurations. Noteworthy features include simultaneous processing of multiple requests and stringent quality checks that maintain model accuracy during optimization. With contributions encouraged from the Strix Halo community, Gufo aims to create a comprehensive resource hub for AI enthusiasts, facilitating the ongoing refinement and expansion of this promising inferencing technology.
Loading comments...
loading comments...