We built the new fastest API for GLM-5.2 (www.baseten.co)

🤖 AI Summary
A month after the release of GLM-5.2, a new fast API has been developed, achieving impressive speeds of up to 280 tokens per second and averaging 100 tokens per second. This API has more than doubled the performance of the original launch-day version, as confirmed by benchmarks from Artificial Analysis. The enhancements are a result of targeted optimizations, including an improved scheduler, upgraded NVFP4 weights, and a refined speculative decoding profile, all of which significantly enhance both latency and throughput. The fast API, specifically designed for low-latency applications such as coding and agent tasks, utilizes Tensor and Expert Parallelism while minimizing max batch sizes to reduce resource competition. Although this optimization comes with a 50% increase in input and output token prices, the feedback from users highlights notable performance improvements in practical use cases, transcending benchmark results. The continued investment in optimizations, including an upcoming update to the speculative decoding algorithm, signifies a strong commitment to enhancing GLM-5.2's capabilities and user experience. The new API is now publicly accessible, inviting users to leverage this advanced tool at Baseten.
Loading comments...
loading comments...