Reflection Beam vs GLM-5.3 and Kimi K3: benchmarks for the new 501B open model
Back to the blog
Articleinteligência artificialtecnologiainovação

Reflection Beam vs GLM-5.3 and Kimi K3: benchmarks for the new 501B open model

MafraOctober 6, 20265 min read

Reflection AI launched Beam on 10/05/2026 (October 5): an open-weights model with 501 billion parameters, 23 billion active per token, an Apache 2.0 license and a focus on code and agents. The question for anyone who programs is direct: does it beat GLM-5.3 or Kimi K3? By Reflection's own table, no. It beats GLM-5.2 on 3 of 7 tests, and loses to GLM-5.3 on all 8 tests where both have a score.

That does not make Beam irrelevant. Its pitch is different: scores close to GLM-5.2 while using 3 to 4 times less generation compute, according to Reflection. This article separates what has been measured from what is still a promise.

Does Reflection Beam beat GLM-5.3 and Kimi K3?

No. In the table Reflection published, Beam trails GLM-5.3 on all 8 tests where both have a score, and trails Kimi K3 on every test where Kimi appears. The gap ranges from 1.2 points (GPQA Diamond) to 26.4 points (SWE Atlas Codebase QnA).

BenchmarkBeamGLM-5.2GLM-5.3Kimi K3
DeepSWE v1.144.444.061.068.0
SWE-Bench Pro v165.562.1no scoreno score
SWE-Bench Pro v2-Hard77.2no score84.388.2
Terminal Bench v2.180.181.088.288.3
SWE Atlas Codebase QnA34.6no score61.068.0
MCP Atlas78.777.884.282.3
SciCode49.7no score59.058.7
HLE without tools36.240.542.346.9
GPQA Diamond90.591.291.793.5

Source: Reflection AI, "Introducing Beam", reflection.ai/blog/introducing-beam, accessed 10/06/2026. All scores are reported by Reflection itself. "No score" means Reflection did not publish the number.

Against GLM-5.2, does Beam tie or lose?

It is roughly a tie: Beam wins 3 tests and loses 4. GLM-5.2 is the model Reflection uses as its efficiency yardstick.

BenchmarkBeamGLM-5.2Winner
DeepSWE v1.144.444.0Beam (+0.4)
SWE-Bench Pro v165.562.1Beam (+3.4)
MCP Atlas78.777.8Beam (+0.9)
Terminal Bench v2.180.181.0GLM-5.2 (+0.9)
AIME 202697.899.2GLM-5.2 (+1.4)
GPQA Diamond90.591.2GLM-5.2 (+0.7)
HLE without tools36.240.5GLM-5.2 (+4.3)

Beam's wins are under 1 point on two tests and 3.4 points on one. Its heaviest loss is on HLE, 4.3 points. On agentic coding the balance is small in both directions.

Where is Beam weak, and where is its promise?

The weak spot is understanding a large repository: on SWE Atlas Codebase QnA it scores 34.6, against 61.0 for GLM-5.3 and 68.0 for Kimi K3. The promise is efficiency, which Reflection says is 3 to 4 times less generation compute than GLM-5.2.

That second part is a vendor claim, with no independent measurement so far. Data you need to decide is also missing:

What drives the choiceBeam's status on 10/06/2026
Price per million tokensNot published at launch
Open weightsPromised for "later this month" (October 2026), not yet released
LicenseApache 2.0 (applies once the weights are out)
AccessWaitlist at platform.reflection.ai, in beta
Context1 million tokens in training, per Reflection; the beta API has a lower limit
Input and outputText only, no image, audio or file
Independent evaluationNone yet

Without a price, there is no cost per benchmark point. For GLM-5.2, Artificial Analysis lists US$ 1.40 per million input tokens and US$ 4.40 per million output tokens (checked 10/06/2026, on a page that marks it as deprecated in favor of GLM-5.3). Beam's price stays open until Reflection publishes it.

How do you test Beam on your own code?

Request access at platform.reflection.ai and point any OpenAI-compatible client at Reflection's API. According to Kingy AI's analysis of the documentation, the base URL is https://api.reflection.ai/openai/v1 and the model identifier is Beam-501B-A23B.

curl https://api.reflection.ai/openai/v1/chat/completions \
  -H "Authorization: Bearer $REFLECTION_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Beam-501B-A23B",
    "messages": [{"role": "user", "content": "Explain what this diff changes and list the risks."}]
  }'

Run the same task on two models and count how many attempts each one needed. A benchmark score does not replace that. The script below summarizes a scoreboard comparison using the scores from the table above, and accepts yours.

# Scoreboard: on how many tests model A beats B (scores from Reflection's table, 10/05/2026)
beam = {"DeepSWE": 44.4, "SWE Pro v1": 65.5, "Terminal Bench": 80.1,
        "MCP Atlas": 78.7, "HLE": 36.2, "GPQA": 90.5}
glm52 = {"DeepSWE": 44.0, "SWE Pro v1": 62.1, "Terminal Bench": 81.0,
         "MCP Atlas": 77.8, "HLE": 40.5, "GPQA": 91.2}

common = beam.keys() & glm52.keys()
wins = [t for t in sorted(common) if beam[t] > glm52[t]]
print(f"Beam wins {len(wins)} of {len(common)}: {wins}")

So which open model should you pick for coding today?

For production work today, GLM-5.3 or Kimi K3, because they score higher and are already available. Beam is a candidate to watch once it has a price, weights and independent measurement.

SituationPickWhy
Hard agentic coding, large repositoryKimi K3 or GLM-5.3Ahead of Beam on every coding test in the table, by a wide margin on Atlas QnA and DeepSWE
You want the best open model available nowGLM-5.3Already available and ahead of Beam on all 8 comparable tests
You want to run it on your own GPU once weights shipKeep watching BeamApache 2.0 and 23B active are attractive, but weights, price and third-party measurement are missing
Deciding by benchmark score aloneDon'tIt measures one slice. Run your own task on both

When the gap between two models is a few points, what decides is how many attempts you can afford to make. On Verboo Code the plans run with unlimited tokens, so running the same task on more than one model with /model to compare does not raise your bill.

Read also

Enjoyed this article?
Share knowledge with your network.
// Read also

Related articles