VibeSys Builds a Qwen3.5-397B Engine on MI300A, 2.3× Tuned SGLang (syfi.cs.washington.edu)

🤖 AI Summary
VibeSys has successfully developed a high-performance serving engine for the Qwen3.5-397B-A17B model using a multi-agent system, achieving a remarkable throughput of 2,242 tokens per second (tok/s) on AMD MI300A hardware. This 2.33x increase in goodput over the baseline shows the potential of automating engine design without human intervention. VibeSys agents optimized the engine for a multi-turn chat workload, effectively tuning the model to achieve exceptional performance over a 105-hour period while leveraging mixed attention mechanisms and specialized optimizations for model weights. The significance of this achievement lies in its demonstration of how AI-driven automation can enhance the deployment and performance of large-scale machine learning models, particularly in optimizing serving engines. Key technical highlights include the development of a custom HIP mixture-of-experts (MoE) kernel, effective prefix caching, and the replacement of conventional collective communications with more efficient methods. The optimization pipeline utilized a combination of parallel processing and innovative graph capturing techniques, leading to substantial gains in throughput and responsiveness. This work underscores the evolving capabilities of AI in system design and opens avenues for further advancements in scaling AI deployments efficiently.
Loading comments...
loading comments...