The Allen Institute for AI has posted Olmo-core 3 on Hugging Face, a redesigned open training system for mixture-of-experts models. An MoE can hold many parameters while each input uses only some of them. The full model still has to sit in GPU memory and be updated during training. Sending each input to the right experts across a cluster has its own communication cost. As the model grows, that cost can erase the savings from activating only part of it.

[1]
Tempera: a full grid of pale squares with only four filled rust red.
The grid fills the sheet. Only four squares are rust red, and they do not sit together. That is a large expert pool that still activates a few experts per token. An illustration, not a training dashboard., AI-generated illustration, not a news photograph

In one benchmark, the expert pool grew from 8 to 128 while each token still selected only four experts, and active parameters per token stayed near 3.2 billion. Total parameters rose from 4.6 billion to 47 billion, and training throughput fell by less than 5 percent. The same infrastructure has been benchmarked above one trillion total parameters.

The earlier MoE path in Olmo-core used fully sharded data parallelism and gathered weights again for every small batch. Olmo-core 3 uses distributed data parallelism instead: experts stay on the GPUs, and the relevant data is routed to them. In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE reached 52,000 tokens per second per GPU on the new stack, against 19,400 on the earlier implementation, about 2.7 times the throughput. That test compares Ai2’s new stack with its own old one, not with Megatron-Core.

[1]

Three splits decide what each GPU stores. Expert parallelism keeps only part of the expert pool on each GPU. Pipeline parallelism splits layers across groups of GPUs. A distributed optimizer spreads optimizer state instead of copying all of it onto every GPU. On the routing side, data is placed directly into expert input buffers, routing metadata stays on the GPU, and grouped matrix multiplies combine many small expert computations.

MXFP8 stores some values in fewer bits. In a controlled run on four B300 GPUs, with work spread evenly across experts, turning MXFP8 on where it helped most raised end-to-end training throughput by about 21 percent versus BF16. Peak active memory fell from 103 GiB to 95 GiB. Most of the gain came from the feed-forward computation and from moving data between experts, not from attention alone. The post says speeding up one part of training can add cost elsewhere, and moving fewer bits does not help if converting the format takes too long.

[1]

On B300 GPUs they benchmarked a 1.2-trillion-parameter model with 58.36 billion parameters active per token, across 512 GPUs. The highest throughput they observed was 858 trillion floating-point operations per second per GPU. Those runs used random routing to measure the system, not the quality of a trained model. A short capacity test with DeepEP v2 reached 2.38 trillion total parameters. It was not a full training run. It shows the scale the stack can reach.

The post also records results that are not gains. A score meant to encourage balanced routing could improve even while the real workload grew less balanced. They call that token gerrymandering. Lowering experts’ learning rates because they see fewer tokens did not help in the model family they tested. With the same matrix shape, different input values changed how long the GPU calculation took. Overlapping communication and computation on separate GPU streams sometimes made the whole step slower.

The post points to a technical report and to GitHub. The figures above are from the blog, not a recheck of the report’s tables. It says the Olmo they are building after this one will use an MoE, and that they want it to be their most capable Olmo, on their largest dataset and longest context. That is the aim, not a model this release has finished training.

[1]

要点

  • The expert pool grew from 8 to 128, with four experts per token and about 3.2 billion active parameters. Total parameters rose from 4.6 billion to 47 billion, and throughput fell by less than 5 percent.
  • On eight B300 GPUs the new stack reached 52,000 tokens per second per GPU, versus 19,400 before, about 2.7 times. The comparison is with Ai2’s own earlier stack.
  • On four B300 GPUs, MXFP8 raised throughput about 21 percent versus BF16, and peak active memory fell from 103 GiB to 95 GiB.
  • The highest observed rate on a 1.2-trillion-parameter, 512-GPU run was 858 TFLOP/s per GPU, using random routing, not a quality score. The 2.38-trillion figure is a short capacity test.