Reflection AI launched Beam on 10/05/2026 (October 5): an open-weights model with 501 billion parameters, 23 billion active per token, an Apache 2.0 license and a focus on code and agents. The question for anyone who programs is direct: does it beat GLM-5.3 or Kimi K3? By Reflection's own table, no. It beats GLM-5.2 on 3 of 7 tests, and loses to GLM-5.3 on all 8 tests where both have a score.
That does not make Beam irrelevant. Its pitch is different: scores close to GLM-5.2 while using 3 to 4 times less generation compute, according to Reflection. This article separates what has been measured from what is still a promise.
Does Reflection Beam beat GLM-5.3 and Kimi K3?
No. In the table Reflection published, Beam trails GLM-5.3 on all 8 tests where both have a score, and trails Kimi K3 on every test where Kimi appears. The gap ranges from 1.2 points (GPQA Diamond) to 26.4 points (SWE Atlas Codebase QnA).
| Benchmark | Beam | GLM-5.2 | GLM-5.3 | Kimi K3 |
|---|---|---|---|---|
| DeepSWE v1.1 | 44.4 | 44.0 | 61.0 | 68.0 |
| SWE-Bench Pro v1 | 65.5 | 62.1 | no score | no score |
| SWE-Bench Pro v2-Hard | 77.2 | no score | 84.3 | 88.2 |
| Terminal Bench v2.1 | 80.1 | 81.0 | 88.2 | 88.3 |
| SWE Atlas Codebase QnA | 34.6 | no score | 61.0 | 68.0 |
| MCP Atlas | 78.7 | 77.8 | 84.2 | 82.3 |
| SciCode | 49.7 | no score | 59.0 | 58.7 |
| HLE without tools | 36.2 | 40.5 | 42.3 | 46.9 |
| GPQA Diamond | 90.5 | 91.2 | 91.7 | 93.5 |
Source: Reflection AI, "Introducing Beam", reflection.ai/blog/introducing-beam, accessed 10/06/2026. All scores are reported by Reflection itself. "No score" means Reflection did not publish the number.
Against GLM-5.2, does Beam tie or lose?
It is roughly a tie: Beam wins 3 tests and loses 4. GLM-5.2 is the model Reflection uses as its efficiency yardstick.
| Benchmark | Beam | GLM-5.2 | Winner |
|---|---|---|---|
| DeepSWE v1.1 | 44.4 | 44.0 | Beam (+0.4) |
| SWE-Bench Pro v1 | 65.5 | 62.1 | Beam (+3.4) |
| MCP Atlas | 78.7 | 77.8 | Beam (+0.9) |
| Terminal Bench v2.1 | 80.1 | 81.0 | GLM-5.2 (+0.9) |
| AIME 2026 | 97.8 | 99.2 | GLM-5.2 (+1.4) |
| GPQA Diamond | 90.5 | 91.2 | GLM-5.2 (+0.7) |
| HLE without tools | 36.2 | 40.5 | GLM-5.2 (+4.3) |
Beam's wins are under 1 point on two tests and 3.4 points on one. Its heaviest loss is on HLE, 4.3 points. On agentic coding the balance is small in both directions.
Where is Beam weak, and where is its promise?
The weak spot is understanding a large repository: on SWE Atlas Codebase QnA it scores 34.6, against 61.0 for GLM-5.3 and 68.0 for Kimi K3. The promise is efficiency, which Reflection says is 3 to 4 times less generation compute than GLM-5.2.
That second part is a vendor claim, with no independent measurement so far. Data you need to decide is also missing:
| What drives the choice | Beam's status on 10/06/2026 |
|---|---|
| Price per million tokens | Not published at launch |
| Open weights | Promised for "later this month" (October 2026), not yet released |
| License | Apache 2.0 (applies once the weights are out) |
| Access | Waitlist at platform.reflection.ai, in beta |
| Context | 1 million tokens in training, per Reflection; the beta API has a lower limit |
| Input and output | Text only, no image, audio or file |
| Independent evaluation | None yet |
Without a price, there is no cost per benchmark point. For GLM-5.2, Artificial Analysis lists US$ 1.40 per million input tokens and US$ 4.40 per million output tokens (checked 10/06/2026, on a page that marks it as deprecated in favor of GLM-5.3). Beam's price stays open until Reflection publishes it.
How do you test Beam on your own code?
Request access at platform.reflection.ai and point any OpenAI-compatible client at Reflection's API. According to Kingy AI's analysis of the documentation, the base URL is https://api.reflection.ai/openai/v1 and the model identifier is Beam-501B-A23B.
curl https://api.reflection.ai/openai/v1/chat/completions \
-H "Authorization: Bearer $REFLECTION_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "Beam-501B-A23B",
"messages": [{"role": "user", "content": "Explain what this diff changes and list the risks."}]
}'
Run the same task on two models and count how many attempts each one needed. A benchmark score does not replace that. The script below summarizes a scoreboard comparison using the scores from the table above, and accepts yours.
# Scoreboard: on how many tests model A beats B (scores from Reflection's table, 10/05/2026)
beam = {"DeepSWE": 44.4, "SWE Pro v1": 65.5, "Terminal Bench": 80.1,
"MCP Atlas": 78.7, "HLE": 36.2, "GPQA": 90.5}
glm52 = {"DeepSWE": 44.0, "SWE Pro v1": 62.1, "Terminal Bench": 81.0,
"MCP Atlas": 77.8, "HLE": 40.5, "GPQA": 91.2}
common = beam.keys() & glm52.keys()
wins = [t for t in sorted(common) if beam[t] > glm52[t]]
print(f"Beam wins {len(wins)} of {len(common)}: {wins}")
So which open model should you pick for coding today?
For production work today, GLM-5.3 or Kimi K3, because they score higher and are already available. Beam is a candidate to watch once it has a price, weights and independent measurement.
| Situation | Pick | Why |
|---|---|---|
| Hard agentic coding, large repository | Kimi K3 or GLM-5.3 | Ahead of Beam on every coding test in the table, by a wide margin on Atlas QnA and DeepSWE |
| You want the best open model available now | GLM-5.3 | Already available and ahead of Beam on all 8 comparable tests |
| You want to run it on your own GPU once weights ship | Keep watching Beam | Apache 2.0 and 23B active are attractive, but weights, price and third-party measurement are missing |
| Deciding by benchmark score alone | Don't | It measures one slice. Run your own task on both |
When the gap between two models is a few points, what decides is how many attempts you can afford to make. On Verboo Code the plans run with unlimited tokens, so running the same task on more than one model with /model to compare does not raise your bill.



