Reflection Unveils 501B-Parameter Beam
- •Reflection announced Beam, its first open-weight model, on October 5, 2026.
- •Beam has 501 billion total parameters and uses 23 billion per token; training used 10,500 GB300 GPUs for four weeks.
- •Reflection plans to release Beam’s weights under Apache 2.0 in October, with FP8 and NVFP4 quantized versions also planned.
U.S. AI startup Reflection announced Beam on October 5, 2026, its first open-weight model for coding, reasoning, and AI-agent tasks. The model has 501 billion total parameters, with 23 billion active for each text-processing step. Reflection has given early access to some users and plans to release its trained parameters, known as weights, under the Apache 2.0 license in October.
Beam uses a Mixture-of-Experts architecture, which selects specialist components based on the input, and activates 23 billion rather than all 501 billion parameters for each token. Reflection pretrained the model from scratch, then used reinforcement learning to improve its reasoning and tool-use abilities. Its reported scores were 44.4 on DeepSWE v1.1, which measures extended software development tasks, and 80.1 on Terminal Bench v2.1, which measures terminal-based tasks. GLM 5.2 scored 44.0 and 81.0, respectively; Qwen 3.8-Max scored 51.0 and 86.6, beating Beam on both evaluations.
Reflection said Beam matched GLM 5.2 on advanced reasoning evaluations using about one-third to one-quarter as much inference compute. The estimate was based on the number of parameters used and the length of the reasoning process and final answer; it excludes input processing, computation for long contexts, and service-operation overhead. It is not a direct comparison of user prices or response speed. Users can adjust the model’s inference compute for each task: lower settings produce shorter reasoning, while higher settings allow longer reasoning for difficult tasks.
Pretraining used 23.8 trillion tokens from selected web material, public documents, and licensed data, covering code, technical documents, mathematics, and science. For four weeks of reinforcement learning, Reflection ran 10,500 NVIDIA GB300 GPUs and generated more than 100 million task attempts. The company prepared about 1 million training environments for coding, tool use, and science and technology tasks, repeatedly adjusting their difficulty and quality. Reflection said DeepSWE evaluation performance continued to improve as it increased reinforcement-learning compute.
Beam is undergoing red-teaming to investigate misuse and unexpected behavior, along with final evaluations before release. Reflection plans to publish the weights, a technical report, a model card covering performance and limitations, and development materials for running, evaluating, and further training the model in October. The weights will use the Apache 2.0 license. Reflection also plans to offer FP8 and NVFP4 quantized versions to reduce memory use and computational load, and is accepting early-access registrations.