Ai2 has released Olmo-core 3, a major redesign of the open training framework behind its Olmo program. The focus is mixture-of-experts training: keeping large expert pools distributed across GPUs, routing tokens efficiently and avoiding communication patterns that erase the theoretical compute advantage of sparse models. Ai2 says the stack is one of the core systems behind its next generation of Olmo and is designed to scale into the trillion-parameter range.
The engineering shift is from an earlier fully sharded approach toward a system built around distributed data parallelism, expert parallelism, pipeline parallelism and a distributed optimizer. Ai2 reports that, in a preliminary eight-B300 test, a 47B-parameter MoE processed about 52,000 tokens per second per GPU compared with 19,400 for its earlier implementation, roughly a 2.7× increase. In a separate controlled test, enabling MXFP8 where it helped most raised end-to-end throughput by about 21% while reducing peak active memory. These are publisher-reported systems measurements, not model-quality results.
The broader importance is provenance. Open weights are much more informative when the training code, parallelism choices and failed experiments are inspectable too. Ai2 documents not only the optimizations it kept but also cases where apparently promising techniques did not improve end-to-end performance. For Open Model Weights, this is the kind of upstream evidence that can eventually make lineage and training-method records richer without guessing: when a future Olmo release cites Olmo-core 3, the model record can point to an open implementation rather than merely a marketing description.
What this changes for the evidence layer
Training-stack news, not a model-weight release. Future Olmo model records can link back to this stack when the publisher identifies it as part of the training provenance.
Ai2 / Hugging Face
This brief is based on the cited primary source. Performance, benchmark and comparative claims remain attributed unless Open Model Weights publishes an independent measurement.
Open source article ↗