Instella-Moe: An Open Mixture-of-Experts Language Model (rocm.blogs.amd.com)

🤖 AI Summary
AMD has unveiled Instella-MoE, a cutting-edge fully open Mixture-of-Experts (MoE) language model with a total of 16 billion parameters, including 2.8 billion active parameters per token. Built from scratch on AMD's MI300X and MI325X GPUs using the ROCm software stack, this model integrates innovative architectural features such as Gated Multi-head Latent Attention and FarSkip-Collective connectivity. Performance benchmarks demonstrate that Instella-MoE stands strong against both traditional dense models and other MoE alternatives, positioning it as a leading open language model at its scale. This release is significant for the AI and machine learning community as it furthers AMD's mission of promoting open-source research and collaboration. By making all training artifacts—including model weights, configurations, and code—publicly available, AMD aims to enhance reproducibility and foster innovation within the field. The model's training pipeline consists of multiple stages, from pre-training on a diverse dataset to fine-tuning for instruction-following and reasoning capabilities. Notably, the incorporation of advanced techniques like communication-computation overlap during training enhances efficiency, yielding quicker inference times, which is crucial for real-world applications.
Loading comments...
loading comments...