GLM 5.3 flash on AMD GPUs 670 tok/s (twitter.com)

🤖 AI Summary
RunInfra has announced a significant update to its GLM 5.3 Flash model, now optimized for AMD GPUs. This major release improves performance to an impressive 670 tokens per second on the Vercel AI Gateway. It comes with a competitive pricing structure—$0.11 per million input tokens, $0.45 per million output tokens, and $0.03 per million cached tokens—while supporting a context of up to 1 million tokens. The upgraded model retains its original API and features, but now boasts enhanced capacity and speed, alongside a 99.7% cache hit rate, providing users with detailed cache logs for better transparency. This update is significant for the AI/ML community as it broadens the compatibility of high-performance models across different hardware platforms, specifically AMD, which can enable more users to access powerful AI tools. Additionally, GLM 5.3 Flash offers OpenAI-compatible chat completions and Anthropic-compatible messaging. With features such as text and image input, tool calling, JSON mode, and streaming capabilities, it positions itself as a robust solution for developers, all while ensuring zero data retention—meaning user data is never used for training.
Loading comments...
loading comments...