Reflection has announced Beam, a large language model aimed at software engineering and tool-using agents, but nobody outside the company can download or run it yet. In a post on its own blog, Reflection says it will publish the weights under an Apache 2.0 license later in October, along with a technical report, a model card, and tooling for running and fine-tuning the model. For now, access is a waitlist for a select group, and the model is still in final red-teaming and evaluation.

That gap shapes how to read everything else. Every score below comes from Reflection’s announcement of its own first open-weight model. For rival models, the post draws numbers from Artificial Analysis and DataCurve, so no outsider has tested Beam itself.

The two parameter counts are the part people trip over. Beam holds 501 billion parameters in total, the adjustable numbers that store what a model has learned. It is a mixture-of-experts design: the network is divided into many specialist sub-networks, and a router sends each piece of text to only a few of them. So roughly 23 billion parameters work on any single token while the rest sit idle for that step. Running cost tracks the 23 billion, which is Reflection’s case for calling the model efficient.

The training figures are the company’s own disclosures and cannot be checked. Reflection says it pretrained on 23.8 trillion tokens (chunks of text) from the web and from licensed proprietary datasets. It then ran reinforcement learning, where the model attempts a task, gets graded, and adjusts. Each attempt from start to graded finish is a rollout. Reflection reports more than 100 million of them over four weeks on 10,500 Nvidia GB300 chips, and says Inkling, a model it compares against, used 30 million.

On coding, Reflection’s chosen comparisons give a mixed picture. On SWE-Bench Verified, Beam scores 80.9 against 77.6 for Inkling and 70.7 for Nemotron 3 Ultra. On SWE-Bench Pro v1, it posts 65.5, ahead of GLM 5.2 at 62.1 and behind Qwen 3.8-Max at 67.7. On Terminal-Bench 2.1, Beam reaches 80.1, just under GLM 5.2 at 81.0 and well short of Kimi K3 at 88.3 and Qwen 3.8-Max at 86.6.

The gaps widen on DeepSWE, a test of fixing real software issues. Beam scores 44.4, level with GLM 5.2 at 44.0, but Kimi K3 posts 68.0 and DeepSeek V4.1 Flash 74.2. Reflection concedes that Kimi K3 stays ahead on raw capability. Its pitch is that Beam gets close at far lower compute, and it says Beam matches GLM 5.2 on advanced reasoning tests while using three to four times less inference compute.

That efficiency number deserves care. Reflection estimates compute from active parameters multiplied by the average tokens generated per attempt, and its own note says the figure leaves out prompt processing and serving overhead. It is an approximation, not a measured bill. The post also gives no price, no hosting terms, and no independent evaluation.

One practical catch sits behind the 23 billion figure. A mixture-of-experts model still has to keep all of its experts available in memory, so self-hosting Beam will need hardware sized for 501 billion parameters even though each token uses a small slice. The saving shows up in speed and compute per token, not in the memory footprint.

Teams considering Beam should treat the October release as the real starting line. Run it on your own repositories once the weights land, because a model that trails Kimi K3 and DeepSeek V4.1 Flash on Reflection’s own tests has to win on cost to earn a place.

Reflection, in a post on its own company blog introducing Beam; the post carries no publication date in the scraped copy, so none is given here.