Alibaba Cloud: AI Models, Reducing Footprint of Nvidia GPUs, and Cloud Streaming (boilingsteam.com)

🤖 AI Summary
At TGS 2025 the Alibaba Cloud booth spotlighted the company’s growing AI stack: large teams behind Qwen and Wan models, frequent releases of open weights you can run locally, and a preview of WAN 2.5 that can already generate ~10‑second videos with audio. Staff emphasized the scale and secrecy of their model efforts (hundreds of engineers) and pointed to Alibaba’s willingness to publish weights for broader use, making advanced models more accessible to developers outside of the usual US cloud/SDK ecosystem. The conversation also revealed a strategic hardware/software play: China’s export limits on Nvidia H100/H200 have pushed Alibaba to rely on H20 GPUs where available and to develop an internal 7nm inference chip plus a CUDA‑like compatibility layer (not 100% compatible) for running workloads. Training still favors Nvidia hardware today, but Alibaba’s stack suggests growing capability for inference independence — a development corroborated by SCMP — with potential long‑term implications for model portability, software fragmentation, and the global GPU market. Alibaba Cloud is also packaging low‑latency Linux‑based game streaming, content distribution and server management services, positioning itself as a cheaper AWS alternative in gaming and cloud AI infrastructure.
Loading comments...
loading comments...