The Allen Institute for AI, known as Ai2, has published Olmo-core 3, the open training software it will use to build its next generation of Olmo language models, and its own benchmarks put the code at more than one trillion parameters.
The target is a design called mixture of experts, or MoE. Instead of running every part of a model for every word, an MoE routes each token (a small chunk of text) to a few specialist sub-networks and leaves the rest idle. That keeps the cost per token low while total capacity climbs. Ai2 says the catch appears at scale: the full model must still sit in GPU memory, and shuttling tokens between specialists across a cluster burns time that can eat the savings.
Ai2 supports its claim with a scaling test. It grew the pool of specialists from 8 to 128 while still choosing four per token, so the active parameters stayed near 3.2 billion. Total size jumped from 4.6 billion to 47 billion parameters, yet training throughput dropped by less than 5 percent. Every figure in this piece comes from Ai2’s own blog post, not an outside lab.
The speedup over Ai2’s previous approach is larger. Ai2 calls its test preliminary: eight NVIDIA B300 chips training a 47-billion-parameter MoE. Each chip handled 52,000 tokens every second under the new code. The old implementation managed 19,400 on the same setup, so the gain is roughly 2.7 times. The old design kept re-gathering model weights for every small batch. The new one leaves the specialists parked on their chips and sends the data to them.
A lower-precision number format called MXFP8 adds more. On four B300 chips, Ai2 measured about 21 percent higher throughput than its higher-precision baseline, with peak memory falling from 103 GiB to 95 GiB. That test spread work evenly across specialists, which real training rarely does.
The headline scale numbers carry heavier caveats, and Ai2 states them itself. A 1.2-trillion-parameter configuration spread over 512 GPUs peaked at 858 TFLOP/s per GPU, and each token touched 58.36 billion of its parameters. A separate experiment reached 2.38 trillion parameters using an alternative communication method. Both used random routing, so they measure how fast the machinery runs, not whether a trained model would be any good. The 2.38 trillion run was a short capacity test, not sustained training.
The more useful reading for builders may be the failures Ai2 documents in its technical report. One routing score looked better even as the actual workload got more lopsided, which the team calls token gerrymandering. Lowering the learning rate for specialists that see fewer tokens did not help in the models it tested. Overlapping communication with computation sometimes made training slower. Those are the kinds of dead ends that cost other teams weeks.
The release also has a market angle. Ai2 names NVIDIA’s Megatron-Core as the established option for large MoE training, and says its own stack is an integrated alternative inside the framework behind Olmo. It gives no head-to-head comparison against Megatron-Core in the post, so the only measured comparison is Ai2 against its earlier self.
Ai2 says the next Olmo will be an MoE trained on its largest dataset with its longest context window, but it gives no release date or size. For academic labs and smaller companies priced out of frontier training, the practical test is whether the code reproduces these throughput numbers on hardware they can actually rent. Anyone planning an MoE run this year should treat the GitHub release as the starting point for that benchmark.
Reported by the Allen Institute for AI on its own blog. The source page carries no publication date.