Beam: A 501B Open-Weight Model Built for Coding and Agentic Workloads

Reflection introduces Beam, a 501B-parameter open-weight model with 23B active, designed for coding, reasoning, and agentic tasks with frontier-level inference

What Beam delivers

Reflection has introduced Beam, a 501 billion-parameter open-weight model with 23 billion parameters active per token, designed to advance coding, reasoning, and agentic workloads. Beam is a sparse Mixture-of-Experts (MoE) model that the company built through major investments in pretraining and large-scale reinforcement learning (RL). Final red-teaming and evaluations are in progress, with weights, a technical report, a model card, and developer artifacts scheduled for release later this month.

Reflection pretrained Beam on 23.8 trillion curated, high-quality tokens from the web and proprietary licensed datasets, matching or outperforming similar-sized open base models. The model was then trained with high-compute RL, producing capabilities competitive with leading open-weight models while using significantly less compute at inference.

How Beam compares to other open models

Beam is trained with particular focus on coding and agentic performance. Across published benchmark scores, Beam advances the Western open-weight frontier and is competitive with larger open models like GLM 5.2 while approaching Qwen 3.8-Max on coding and agentic tasks. On frontier open models like Kimi K3 that remain ahead in raw capability, Beam’s advantage is inference-time efficiency.

h3>

Key benchmark results reported include:

  • DeepSWE v1.1: Beam 44.4, GLM 5.2 61.0, Kimi K3 68.0, Qwen 3.8 Max 51.0, DeepSeek V4.1 Flash 74.2.
  • SWE Bench Pro v2-Hard: Beam 77.2, GLM 5.3 84.3, Kimi K3 88.2.
  • SWE Bench Pro v1: Beam 65.5, GLM 5.2 62.1, Qwen 3.8 Max 67.7.
  • Terminal Bench v2.1: Beam 80.1, GLM 5.2 81.0, GLM 5.3 88.2, Kimi K3 88.3, Qwen 3.8 Max 86.6, DeepSeek V4.1 Flash 90.6.
  • SWEBench Verified: Beam 80.9, Inkling 77.6, Nemotron 3 Ultra 70.7.
  • SWEBench Multilingual: Beam 78.0, Nemotron 3 Ultra 67.7.
  • AIME 2026: Beam 97.8, GLM 5.2 99.2, Inkling 97.1.
  • GPQA Diamond: Beam 90.5, GLM 5.2 91.2, Kimi K3 93.5, Qwen 3.8 Max 92.6.
  • MCP Atlas: Beam 78.7, GLM 5.2 77.8, Kimi K3 84.2, Qwen 3.8 Max 84.5.
  • BrowseComp w/ context management: Beam 77.4, Inkling 77.1, Kimi K3 91.2.

On advanced reasoning benchmarks, Beam achieves scores comparable to GLM 5.2 while using 3 to 4 times less inference compute. Efficiency gains are more pronounced against models in the 2T+ parameter family like Qwen 3.8-Max, which require significantly more compute per token. The result is more intelligence per token, positioning Beam as a cost-efficient workhorse for enterprise coding and agentic deployments.

Pretraining and the foundation for reasoning

Beam’s pretraining drew on 23.8 trillion tokens from web sources and proprietary licensed data, with curation designed to match or beat the quality of comparable open base models. The pretraining stage created a stable base that the RL phase could then specialize for reasoning and tool use.

The broader open-model community has been pursuing similar scaling. The DeepSeek-V3 technical report, available on arXiv (2412.19437), describes a 671B-parameter MoE language model with 37B activated per token, pretrained on 14.8 trillion tokens and refined through supervised fine-tuning and reinforcement learning. DeepSeek-V3 demonstrated that MoE architectures can deliver strong performance while requiring only 2.788M H800 GPU hours for full training, without irrecoverable loss spikes or rollbacks. These results established a cost-effective template that models like Beam build on.

High-compute reinforcement learning at scale

Reflection made high-compute RL a central scaling axis. The RL run deployed 10.5K NVIDIA GB300 GPUs for four weeks, generating more than 100 million rollouts with a maximum context length of 256K tokens. Training and grading used approximately 1.3 billion sandboxes. Reflection sourced close to one million high-quality coding, agentic, and STEM environments, primarily through synthetic data pipelines, supplemented by proprietary vendor data and open-source sources.

Across the evaluation suite, capabilities continued to improve as RL compute increased, with no sign of plateau. Reflection describes this as one of the largest-scale RL runs conducted by any open lab.

Beam was trained with asynchronous policy gradients. At scale, policy staleness, where tokens in long rollouts are generated by multiple model checkpoints, becomes a major source of instability. Numerical mismatch between training and inference engines compounds the challenge. Reflection developed new algorithms to maintain stable learning under these conditions, enabling fully asynchronous RL that remains stable even when learning from interactions generated more than a day earlier, up to 107 weight versions behind the current policy.

Learning to reason efficiently

Beam was trained with a controllable length penalty that rewards successful solutions while discouraging unnecessary tokens. Early in RL, performance improved even as completion lengths fell: the model learned to solve tasks more effectively with less reasoning. Later, as agentic capabilities grew, completion lengths increased again, with those additional tokens supporting further performance gains.

Users can control this tradeoff through a reasoning effort parameter: lower settings favor shorter responses, while higher settings allow longer reasoning to improve performance on demanding tasks. This gives flexibility to match reasoning effort to the task and compute budget.

How behavior generalizes with RL

Reflection designed Beam’s RL training to develop reasoning and agentic capabilities that generalize beyond training tasks. During a phase of training on reasoning, software engineering, and terminal tasks, Beam showed consistent gains in browsing despite the absence of browsing tasks from the RL mixture. This transfer suggests Beam was learning broader agentic capabilities that generalize across domains.

When given web access, Beam organically learned to search for and query other large language models, and to use OCR APIs to read documents. Example demonstrations include building a live NYC subway dashboard using public transit data, creating a 3D astronaut free-fall game in p5.js, and preparing a fine-tuning notebook for the Gemma-4 model on a Text2SQL task, an out-of-distribution domain for Beam. The fine-tuning run increased the target model’s accuracy on its held-out test set by 66.5%.

Frontier RL infrastructure

High-compute agentic RL requires generating rollouts, executing tools, evaluating outcomes, and updating the model at scale. Beam’s training sustained an average of 110K concurrent rollouts. Seven infrastructure capabilities made this practical:

  • Fully asynchronous execution: Agents generate rollouts while the trainer learns and publishes new model versions. Each token is tagged with the version that produced it.
  • Flexible compute allocation: Inference-to-training GPU ratios ranged from 3.9:1 to 5.4:1, with the trainer resized across five GPU mesh configurations within the same training lineage.
  • Fast model updates: New weights reached the inference fleet in a median of approximately 12 seconds, using hierarchical distribution across racks over RoCE and locally over NVLink, reducing cross-rack traffic by 75% and making fleet-wide adoption 2.2 times faster.
  • Resilience to inference failures: 71 inference incidents were handled without terminating the training job. Inference capacity recovered in a median of eight minutes, with lost capacity accounting for just 0.02% of elapsed serving GPU-minutes.
  • Environments at scale: The platform supported up to 170K concurrent sandbox environments.

Safety and alignment

Beam is undergoing final red-teaming and evaluations before public release. The broader alignment challenge for large language models in safety-critical domains is addressed in the paper “Deliberative Alignment: Reasoning Enables Safer Language Models” on arXiv (2412.16339), which introduces a paradigm that teaches models safety specifications and trains them to explicitly recall and reason over those specifications before answering. This approach demonstrated improved robustness to jailbreaks and reduced overrefusal rates, with better out-of-distribution generalization. Models like Beam can build on these deliberative alignment techniques during fine-tuning to improve safety adherence.

The path ahead

Beam positions itself as a practical workhorse for enterprise coding and agentic workloads, offering frontier-level capability at a lower inference cost than larger open models. With weights, a technical report, and developer artifacts scheduled for release later this month, developers will be able to evaluate Beam’s efficiency and agentic performance directly. The combination of sparse MoE architecture, 23.8 trillion tokens of pretraining data, and 100 million+ RL rollouts on 10.5K GB300 GPUs represents one of the most ambitious open-weight training efforts to date.

FAQ

What is Beam?

Beam is a 501 billion-parameter open-weight Mixture-of-Experts model with 23 billion active parameters, built by Reflection for coding, reasoning, and agentic workloads. It is designed to deliver frontier-level performance with lower inference compute than larger open models.

How much training compute did Beam use?

Beam’s reinforcement learning run deployed 10.5K NVIDIA GB300 GPUs for four weeks, generating over 100 million rollouts with a maximum context length of 256K tokens and using approximately 1.3 billion sandboxes. The pretraining phase used 23.8 trillion curated tokens.

When will Beam’s weights be available?

Beam is undergoing final red-teaming and evaluations. Reflection plans to release the weights, technical report, model card, and developer artifacts later this month.

Related coverage


This article summarizes reporting from reflection.ai. See our editorial disclaimer for how our articles are produced.

🤖
Is your business visible to AI assistants?

Run a free scan to see your AI Visibility Score, SEO rating, and local citation accuracy.

Check Your Score →