Ai2 Releases Open AI Training Stack With a Reported 2.7× Speed Gain
Olmo-core 3 keeps model experts on GPUs instead of repeatedly moving their weights. Its speed gain comes from a 47-billion-parameter test; separate trillion-parameter experiments measure system scale, not trained-model quality.
Olmo-core 3 changes how mixture-of-experts training work is distributed: instead of repeatedly moving expert weights, it keeps experts on GPUs and routes data to them. Ai2 says this produced about 2.7× the throughput of its earlier implementation in a preliminary test—not a general guarantee—using a 47-billion-parameter model on eight NVIDIA B300 GPUs. The open stack lets researchers modify routing and adapt training to hardware; Ai2 plans to use it for its next-generation Olmo model.
01
Ai2 scaled an expert pool from eight to 128 while selecting four experts per token; capacity reached 47 billion parameters with less than a 5% throughput drop.
02
Selective MXFP8 use increased throughput about 21% over BF16 in a controlled four-GPU test, while peak active memory fell from 103 GiB to 95 GiB.
03
A 1.2-trillion-parameter configuration across 512 B300 GPUs peaked at 858 trillion useful FLOPs per GPU, but used random routing and did not measure model quality.
Ai2 released Olmo-core 3 on October 1, giving researchers an open training framework redesigned for large mixture-of-experts AI models. In a preliminary benchmark, Ai2 measured about 2.7 times the throughput of its earlier implementation. The release targets a specific bottleneck: moving model weights and coordinating work across GPUs can eat into the efficiency these models promise.
This is training infrastructure, not a newly released language model. Ai2 says its next-generation Olmo will use a mixture-of-experts architecture and the new stack. Researchers and developers can also use the open software to train their own models, modify routing decisions and adapt the system to different hardware.
Move the text, not the expert weights
A mixture-of-experts model, or MoE, contains specialized components called experts. Each token—a small unit of text—uses only some of them. That lets a model contain more learned parameters without activating all of them for every input. But the full model still needs GPU memory, training updates and a way to send inputs to the selected experts.
Ai2’s earlier implementation repeatedly gathered and redistributed model weights for small batches of training data. Olmo-core 3 instead keeps experts resident on GPUs and sends the relevant data to them. It divides experts, model layers and the extra information needed for training updates across devices, avoiding a full copy of everything on each GPU.
The preliminary speed comparison used a 47-billion-parameter MoE on eight NVIDIA B300 GPUs. Separately, Ai2 increased an expert pool from eight to 128 while still choosing four experts per token. Total capacity rose from 4.6 billion to 47 billion parameters, while throughput fell by less than 5%. Active parameters stayed roughly fixed at 3.2 billion per token.
Smaller numbers, less traffic
The redesign also addresses smaller sources of overhead. It places routed data directly into expert input buffers, keeps routing information on GPUs rather than copying it back to the CPU, and groups small expert calculations into larger operations that GPUs can execute more efficiently.
Olmo-core 3 supports MXFP8, a lower-precision number format that uses fewer bits for some values. In a controlled test on four B300 GPUs, with work spread uniformly across experts, selective use raised throughput about 21% over the higher-precision BF16 baseline. Peak active memory dropped from 103 GiB to 95 GiB.
Most of that gain came from expert computation and data movement, rather than attention alone. The tradeoff is conversion: moving fewer bits helps only if changing number formats does not consume the savings. Ai2’s design therefore treats computation and communication as connected costs, not separate opportunities to optimize.
Trillion-parameter scale is not a quality score
Ai2 also benchmarked a 1.2-trillion-parameter configuration across 512 B300 GPUs, with 58.36 billion parameters active per token. Its highest observed throughput was 858 trillion floating-point operations per second per GPU, measuring useful model computation. Those tests used random routing to assess system performance, not the quality of a trained model.
An experiment using DeepEP v2, an alternative system for communication between experts, reached 2.38 trillion parameters. That was a short capacity test, not a full training run. It demonstrates a configuration the stack can reach, rather than sustained training performance. Neither trillion-scale result is the test behind the reported 2.7× speed gain.
Ai2 also describes failed approaches and benchmarking pitfalls in its technical report. They show why a promising individual measurement may not translate into a faster or better training process:
A score meant to encourage balanced routing improved even as workloads became less balanced—a failure Ai2 calls token gerrymandering.
Lowering experts’ learning rates because they processed fewer tokens did not improve results in the tested model family.
GPU calculation times changed with input values, even at identical matrix dimensions. Matching shapes alone was not enough for performance comparisons.
Running communication and computation simultaneously on separate GPU streams sometimes slowed the full process instead of speeding it up.
Sources
allenai.orgIntroducing Olmo-core 3: Open, scalable training infrastructure for large MoEs | Ai2
Reader comments
Newest comments first. Replies stay oldest first.