Reflection Introduces Beam With Early Access Ahead of October Weight Release
The 501-billion-parameter model targets coding and tool-using work. Its efficiency claim is an estimate, not a measured deployment-cost saving.
Loading page…
The 501-billion-parameter model targets coding and tool-using work. Its efficiency claim is an estimate, not a measured deployment-cost saving.
Listen to this story
Reflection AI says Beam matches GLM-5.2 on advanced reasoning with three to four times less inference compute, but its comparison uses active parameters and generated-token counts rather than full deployment costs, and TechCrunch reports the claims have not been independently verified. The model remains in red-teaming and evaluation: downloadable weights, a technical report, model card and developer materials are promised for later in October. The pitch is a more compute-efficient option, not a performance lead over every rival.
Beam has 501 billion total parameters but activates 23 billion per token; Reflection says it was pretrained on 23.8 trillion tokens.
Reinforcement learning produced more than 100 million attempted solutions over four weeks using 10,500 Nvidia GB300 GPUs.
Reflection built nearly one million mostly synthetic training environments and says improving task quality helped overcome capability plateaus.
Reflection AI introduced its first open-weight model, Beam, on October 5, opening early-access registration ahead of a promised weight release later this month. The company is pitching a text-only system for coding, reasoning and tasks that involve using tools. Its central claim is efficiency: competitive reasoning results without the same computing demands as some larger open models.
Beam has 501 billion parameters—the learned mathematical values inside the model—but activates 23 billion for each token it processes. Its mixture-of-experts design routes tokens through selected parts of the network rather than using the whole model at once. Reflection’s compute comparison therefore counts active parameters, not the much larger total.
Reflection says it pretrained Beam on 23.8 trillion tokens from curated web material and licensed proprietary datasets. That initial training was followed by reinforcement learning, which trains a model using rewards.
Reflection also trained Beam to avoid unnecessary output. A controllable length penalty rewards successful solutions while discouraging extra tokens. Early in reinforcement learning, scores improved as responses shortened. Later, responses grew longer again as the model developed stronger tool-using behavior, accompanied by further performance gains.
Users can adjust a reasoning-effort setting: lower settings favor shorter responses, while higher settings allow more reasoning for difficult tasks. Reflection says its reinforcement-learning run generated more than 100 million attempted solutions over four weeks on 10,500 Nvidia GB300 GPUs.
The company reports that poor task quality caused capability plateaus and other training problems. Improving those tasks was essential to keeping gains going—not simply adding more computing power.
Reflection says Beam achieves advanced reasoning scores comparable to GLM-5.2 while using three to four times less inference compute—the computation needed to generate answers. TechCrunch reports that the performance claims have not been independently verified.
The company estimates compute using active parameter counts and the average number of generated tokens, including both reasoning and final answers. It draws other models’ evaluation results from Artificial Analysis and DataCurve. That method gives a narrower comparison than a timed, end-to-end deployment test.
Reflection acknowledges that Kimi K3 remains ahead on raw capability. Its published table puts Beam at 80.1 on Terminal-Bench v2.1, versus 88.3 for Kimi K3, and 44.4 on DeepSWE v1.1, versus 68.0. The pitch is a better capability-to-compute balance, not an across-the-board performance lead.
Beam is still undergoing final red-teaming—testing for harmful behavior—and evaluations. Reflection says the October release will include the weights, technical report, model card and developer artifacts. Early-access registration is therefore not the same as a completed downloadable release.
The planned rollout includes distribution through large cloud providers and specialist GPU clouds, plus integrations with open-source libraries. Reflection calls Beam a workhorse for enterprises, the public sector and developers. Those intended users are being offered an efficiency-focused model whose files and full technical materials are still to come.
Loading discussion...
Join the conversation
Explain which tasks would make that trade worthwhile.
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.