Olmo-core 3: Open, scalable training infrastructure for large MoEs (allenai.org)

🤖 AI Summary
Today, the AI research organization Ai2 announced the release of Olmo-core 3, a redesigned framework for training large language models with a focus on a scalable open mixture-of-experts (MoE) architecture. This upgrade aims to facilitate MoE training that can handle trillion-parameter models while maintaining computational efficiency. The new framework supports a revolutionary shift in how models utilize GPU resources, helping academia and smaller labs access advanced model training without prohibitive costs. For instance, benchmarks showed that the Olmo-core 3 can handle 47 billion parameters with a modest drop in training throughput, showcasing its efficiency with routing data to specific experts. Significantly, Olmo-core 3 incorporates advanced techniques like expert parallelism, pipeline parallelism, and a distributed optimizer, which optimize the distribution of the model across GPUs and reduce memory requirements. The new stack resulted in up to 2.7 times higher throughput compared to previous implementations, allowing researchers to experiment with MoEs at unprecedented scales. By emphasizing open-source development, Ai2 encourages researchers to utilize and adapt Olmo-core 3 for their own AI models, reinforcing the organization’s commitment to shared progress in AI technology.
Loading comments...
loading comments...