Nemotron-H: A Family of Accurate, Efficient Hybrid Mamba-Transformer Models (research.nvidia.com)

🤖 AI Summary
The introduction of the Nemotron-H family of hybrid Mamba-Transformer models marks a significant advancement in AI model architecture, featuring versions with 8 billion, 47 billion, and 56 billion parameters. Designed to optimize inference efficiency without compromising accuracy, Nemotron-H models improve inference speed by up to 3x compared to similar-sized state-of-the-art Transformer models. This efficiency is critical as the AI landscape increasingly prioritizes speed and reasoning capabilities. The architecture replaces traditional self-attention layers with Mamba layers, maintaining constant memory and computation during token generation, a novel approach aimed at enhancing overall model intelligence. Pre-trained on 20 trillion tokens in FP8 precision, Nemotron-H-56B-Base showcases impressive performance across various benchmarks, rivaling larger models while utilizing fewer resources. The distillation of the 47B model from the 56B maintains similar accuracy levels, further illustrating the models' efficiency. Additionally, the Nemotron-H-56B serves as the backbone for the Cosmos-Reason 1 project, designed for physical AI applications, highlighting the model's versatility. Released on platforms like Hugging Face, these models provide a robust and accessible tool for the AI/ML community, paving the way for future innovations in model training and application.
Loading comments...
loading comments...