Open & local AI

Reflection AI launches Beam, a low-cost open-weight model for reasoning

Reflection AI has unveiled Beam, an open-weight model claiming to match Chinese rivals in reasoning while using significantly less inference compute.

A metallic cube with glowing circuits floating above server racks.
Illustration generated for this article

Reflection AI, a Brooklyn-based startup founded in 2024, has officially launched Beam, its first frontier open-weight artificial intelligence model. The release marks a significant entry into the competitive landscape of large language models, positioning Beam as a cost-effective alternative to both closed-source systems and existing open-weight options from China and the West.

What happened

The company describes Beam as a text-only mixture-of-experts model designed specifically for advanced reasoning, coding, and agentic tasks. According to Reflection, the model delivers performance comparable to leading Chinese open models on complex benchmarks but operates at a fraction of the token cost and inference time compute required by rivals. This announcement confirms earlier reports indicating the startup was nearing a public launch.

Beam features a total of 501 billion parameters, with 23 billion active parameters during inference. It was pretrained on a massive dataset of 23.8 trillion tokens and supports a context window of 1 million tokens. For comparison, Z.ai’s GLM-5.2 model contains roughly 744 billion total parameters with 40 billion active. Reflection claims that while its performance metrics have not yet been independently verified, internal benchmarks show Beam scoring on par with GLM-5.2 and outperforming current leading Western open models.

The startup is positioning Beam directly against major players including Anthropic, OpenAI, Mistral, Meta, and Cohere. Its most immediate domestic competitor may be Inkling, the open model released in July by Thinking Machines Lab. Reflection’s data suggests Beam outscores Inkling on four specific coding tests, though it is important to note that Inkling is multimodal while Beam is text-only. The company intends to release the model’s weights and full technical details later this month, with distribution planned through hyperscalers, neoclouds, and open source libraries.

How it works

Beam utilizes a mixture-of-experts architecture, a design choice that allows the model to activate only a subset of its total parameters for any given input. This sparsity enables the model to maintain a high total parameter count for knowledge retention while keeping the active parameter count low during operation. The result is reduced computational load during inference, which Reflection claims leads to 3-4x less inference compute usage compared to competitors.

The model was trained using high-compute reinforcement learning techniques. This training approach focuses on optimizing the model’s ability to reason through complex problems rather than just predicting the next word in a sequence. By prioritizing reasoning capabilities, the model aims to handle agentic tasks and coding challenges more effectively than traditional pretraining methods alone might allow.

Key details

  • Beam is a 501-billion-parameter model with 23 billion active parameters per inference step.
  • The model was pretrained on 23.8 trillion tokens and supports a 1 million token context window.
  • Reflection claims Beam uses 3-4x less inference compute than rival models while matching their performance on reasoning benchmarks.
  • The startup has raised approximately $4.7 billion, with backers including Nvidia, Sequoia Capital, and Lightspeed Venture Partners.
  • Reflection has secured over $7 billion in deals with SpaceX and Nebius to access Nvidia GB300 chips through 2029.
  • The company is promoting an "AI factory" concept, allowing enterprises and sovereign nations to train customized local AI systems on proprietary data.

Why it matters

For software engineers and technical leaders, the promise of lower inference costs is critical. As AI applications move from experimental prototypes to production environments, the cost of running large models can become prohibitive. A model that offers frontier-level reasoning capabilities at a significantly lower compute cost could make advanced AI features viable for a broader range of products and services. This efficiency gain is particularly relevant for teams building agentic workflows or complex coding assistants where latency and cost are primary constraints.

Furthermore, the launch intensifies the competition between Western and Chinese open-weight models. Developers who have relied on models from Chinese labs for cost-effective performance now have a domestic alternative that claims similar efficiency. This diversification reduces dependency on any single geographic region for foundational AI infrastructure. It also pressures other Western providers to optimize their own models for efficiency, potentially accelerating innovation in sparse architectures and reinforcement learning training methods across the industry.

What you can do

  • Monitor the official release of Beam’s weights and technical documentation later this month to assess compatibility with your current infrastructure.
  • Evaluate the claimed 3-4x reduction in inference compute against your specific workload requirements, keeping in mind that these figures are currently self-reported.
  • Compare Beam’s text-only capabilities against multimodal alternatives like Inkling if your application requires image or audio processing alongside text reasoning.
  • Investigate the "AI factory" concept if your organization handles sensitive proprietary data and requires localized, customized AI systems rather than relying on public APIs.
  • Review the integration options with hyperscalers and neoclouds to determine the most cost-effective deployment strategy for your team.
  • Keep an eye on independent benchmark results as they emerge to validate Reflection’s performance claims against established standards.

Tools from the Bytechap store

$89

DocBento

Self-hosted document management that reads every scan and answers with page citations.

Live demo

Keep reading

All stories