Reflection’s Beam becomes the most capable open-weight model built outside China
AI company Reflection has announced Beam, its first freely available model. Instead of chasing peak performance, Beam aims to deliver strong results with minimal compute, putting it in direct competition with Chinese open-weight models from Deepseek and Qwen.
Reflection built Beam for coding, logical reasoning, and agentic tasks. The mixture-of-experts model activates 23 billion of its 501 billion total parameters per token, keeping compute costs relatively low. On demanding reasoning tasks, Beam matches GLM 5.2 while using three to four times less compute, according to Reflection. The company is going after businesses that use AI for coding and automated workflows but need to keep operating costs in check.Ad On coding and agent benchmarks, Beam also comes close to the much bigger Qwen3.8-Max. Stronger open models like Kimi K3 still beat it on raw performance, the company says. Reflection says it’s already training a successor that aims to close the remaining gap to top-performing open models.Ad Massive reinforcement learning run powered Beam’s abilities Beam gets its capabilities from a combination of standard large-scale pretraining and a particularly compute-heavy reinforcement learning phase. Reflection says it ran 10,500 Nvidia GB300 GPUs for over four weeks during that RL phase, calling it one of the largest training runs any open lab has done. Performance kept improving through the end of the run without hitting a ceiling. Users can also control how thoroughly Beam reasons through a problem. A tunable parameter lets you choose whether the model answers quickly or takes more time to think on harder tasks, trading off compute cost against output quality.Ad Reflection observed what it calls “emergent capabilities” during training. While running an RL mix of reasoning, software engineering, and terminal tasks, the company noticed Beam getting better at web browsing even though no browsing tasks were part of that training mix. With web access, the model independently learned to query other language models and pull documents from external services. Reflection shared several demos, including a live-updating New York City subway map, a small 3D game, and a notebook for fine-tuning another AI model. Beam is text-only, but Reflection says it can process content from other media formats as long as they’re represented as text.Ad For safety and alignment, Reflection trained a second model and merged it with Beam. The guidelines range from hard rules the model must never break to quality standards like factual accuracy and admitting uncertainty, along with a direct, thorough, and proactive response style.