Hugging Face debuts scalable MoE training system Olmo-core 3
Hugging Face has released Olmo-core 3, an upgraded open-source framework for training large language models using mixture-of-experts (MoE) techniques. The new system lets developers scale MoE models up to over one trillion parameters without overwhelming compute costs. In tests, Hugging Face managed to grow model capacity tenfold while barely sacrificing training speed. The new version also switches to a more efficient data-distribution system, boosting throughput by 2.7 times on modern GPUs.
- Olmo-core 3 supports trillion-parameter MoE models
- Redesigned for better efficiency on modern GPUs
- Switches to distributed data parallelism for faster training
- Benchmarked at 2.7× previous throughput
- Aims to cut costs and barriers for large-scale AI work
Sources covering this
In this story
More in AI
OpenAI debuts GPT-6 Astra Ultrafast on Nvidia Blackwell GPUs
OpenAI has launched GPT-6 Astra Ultrafast, now available in its API and to select ChatGPT Work and Codex users.
Audible adds AI guides and interactive audiobook features
Audible is rolling out three new features to make audiobooks more interactive, starting with select titles.
Amazon launches open-source Strands Decider 2B model
Amazon Web Services has released Strands Decider 2B, a free and open-source AI model designed for decision tasks where you pick between…
Cloudflare launches open-source decision model Clef
Cloudflare has released Clef and Clef-flash, two fast, open-source AI decision models, aimed at delivering quick and consistent…