AI News · Models ·

Reflection unveils Beam, a 501-billion-parameter open-weight model

Reflection unveils Beam, a 501-billion-parameter open-weight model

Reflection AI has unveiled Beam, a mixture-of-experts model with 501 billion total parameters and 23 billion active parameters. The company says it targets coding, reasoning and AI agent tasks, with advanced reasoning scores comparable to Z.ai's GLM-5.2. Its weights remain in final testing ahead of a planned October release.

Key points

  • Beam has 501 billion total parameters and 23 billion active parameters.
  • Reflection targets coding, reasoning and AI agent tasks for businesses and developers.
  • The company claims reasoning scores comparable to GLM-5.2 with lower inference computation time.
  • Full weights are planned for October 2026 following final safety testing and evaluation.
  • Independent confirmation of performance, efficiency and deployment costs was not reported.

What happened: Reflection AI has unveiled Beam, an open-weight AI model aimed at business customers and developers working on coding, reasoning and AI agent tasks. Bloomberg reported that the US startup, founded by two former Google DeepMind researchers, says Beam can rival leading models from the US and China while costing less than many competitors. The announcement is not yet a public release: Beam remains in final safety testing and evaluation, with its full weights planned for release in October 2026.

The details: Beam is a mixture-of-experts model with 501 billion total parameters and 23 billion active parameters. Reflection says it achieves advanced reasoning benchmark scores comparable to GLM-5.2, a flagship open model from China’s Z.ai. GIGAZINE reported that the company claims comparable scores using one-third to one-fourth of the inference computation time. Beam also approaches Qwen 3.8-Max on coding and agent tasks, according to the report. Those comparisons come with a qualification: GIGAZINE noted that Beam still trails models such as Kimi K3 in overall capability. Independent confirmation of the performance and efficiency claims was not reported.

Background: Reflection describes Beam as the result of substantial investment in both pre-training and reinforcement learning. According to GIGAZINE, its pre-training used 23.8 trillion tokens drawn from the web and proprietary licensed datasets. Reflection then used 10,500 NVIDIA GB300 chips for four weeks of reinforcement learning, generating more than 100 million trials in which the model produced answers or sequences of actions. The company says the combination supports both competitive performance and efficient inference, the process of running a trained model to produce responses.

Who it affects: Beam is aimed at teams evaluating models for software development and agent-based work. The benchmarks discussed by GIGAZINE cover coding agents, practical tasks in terminal environments, repository-level software problems and reasoning without external tools. That focus makes the announcement relevant to businesses comparing open-weight alternatives for technical workflows. Reflection’s pitch centers on the combination of capability and lower computation requirements, rather than an outright lead on every task. Actual deployment costs and independent results on business workloads were not reported.

What to watch: The next milestone is whether Reflection releases the full weights as planned after completing its final testing. GIGAZINE reported that the company is accepting waiting-list registrations while that work continues. For prospective users, weight availability would make Beam a candidate for closer evaluation, not settle the purchasing decision. Its claimed efficiency advantage and coding performance still need independent testing. Until then, Beam represents a forthcoming option for teams seeking open-weight coding and agent models, rather than a publicly available model ready for a deployment decision.

Our take

Beam could expand the options for teams seeking open-weight coding and agent models. Deployment decisions should wait for weight availability and independent testing of its performance and costs.

Sources