VeloGB10 – A GB10-specific inference engine (github.com)

🤖 AI Summary
NVIDIA has launched the veloGB10, a specialized inference engine tailored specifically for GB10-based systems, including the NVIDIA DGX Spark and compatible OEM machines. Developed using Rust and CUDA, veloGB10 is designed to support select large language models such as the Qwen 3.5/3.6/3.8 series and Tencent Hy3, offering advanced capabilities like hybrid GatedDeltaNet architectures and sparse attention mechanisms. This engine emphasizes performance optimization through tensor-parallel inference, enabling faster processing across one to four interconnected GB10 systems. The significance of veloGB10 lies in its capability to handle substantial workloads efficiently while ensuring optimized model execution on GB10 hardware. With specialized features like native support for NVFP4 and EXL3 weight formats, along with innovative kernel designs for MoE expert models, this inference engine maximizes throughput and minimizes latency during model inference. Key metrics indicate that processing speeds can reach up to 85 tokens per second on four GB10s, enhancing productivity for AI/ML developers working with expansive datasets and complex models. The introduction of veloBenchmark further facilitates performance assessment, making this tool a valuable asset in the AI/ML landscape.
Loading comments...
loading comments...