Open vs closed models: how much each benchmark point costs (GLM-5.3, Kimi K3, Opus 5.5)
Back to the blog
Articlecomparativomodelos de IAbenchmarksopen source

Open vs closed models: how much each benchmark point costs (GLM-5.3, Kimi K3, Opus 5.5)

MafraSeptember 29, 20266 min read

The best open-weight model on the Artificial Analysis Intelligence Index scores 46. The best closed model scores 58. The 12-point gap is real and it shows. The question that decides what you use day to day is a different one: how much does each point cost, and how much does each extra point cost.

We ran the numbers with public Artificial Analysis data from September 29, 2026 for six models: three open-weight (MiMo-V2.6-Pro, GLM-5.3 and Kimi K3) and three closed models from Anthropic (Claude Opus 5.5, Sonnet 5.5 and Fable 5.1). The result: a point on an open model costs 2.3 to 51 times less, and each extra point on a closed model costs close to 7 times the average open-model point.

How much does each benchmark point cost on open and closed models?

On GLM-5.3, US$ 0.045 per point. On Claude Opus 5.5, US$ 0.103. The most efficient closed model costs 2.3 times more per point than an open model of the same generation.

The math is simple: average cost per index task divided by the score. Cost per task already includes input, cache and output at the lab's API price, and every model ran at maximum effort.

ModelWeightsIndex scoreCost per taskCost per pointOutput tokens on the index
MiMo-V2.6-Pro (Xiaomi)open, MIT46US$ 0.13US$ 0.003140M
GLM-5.3 max (Z.ai)open45US$ 2.01US$ 0.045210M
Kimi K3 max (Moonshot)open44US$ 2.00US$ 0.045160M
Claude Opus 5.5 maxclosed58US$ 5.98US$ 0.103260M
Claude Sonnet 5.5 maxclosed56US$ 7.60US$ 0.136410M
Claude Fable 5.1 maxclosed53US$ 7.63US$ 0.144190M

Source: Artificial Analysis Intelligence Index v4.3.2, individual model pages on artificialanalysis.ai, checked on September 29, 2026. Cost per point calculated by us: cost per task ÷ score.

Why does Sonnet 5.5 cost more per task than Opus 5.5?

Because it spends far more tokens. Sonnet 5.5 charges half of Opus 5.5 per token (US$ 2 and US$ 10 per million, versus US$ 4 and US$ 20), but it generated 410 million output tokens to run the index, versus 260 million for Opus.

The result: Sonnet costs US$ 7.60 per task and Opus US$ 5.98, with a higher score. The same goes for Fable 5.1: more expensive per task and 5 points below Opus 5.5. At maximum effort, both are dominated, meaning there is another model that is both better and cheaper.

The lesson applies to any comparison: price per token is not price per task. A verbose model cancels out a list-price discount.

How much does an extra point on a closed model cost?

Moving from GLM-5.3 to Claude Opus 5.5 adds 13 points and costs US$ 3.97 more per task. Each extra point comes out at US$ 0.31, close to 7 times the average cost of a GLM-5.3 point.

SwitchExtra pointsExtra cost per taskCost per extra point
GLM-5.3 → Claude Opus 5.5+13US$ 3.97US$ 0.31
MiMo-V2.6-Pro → Claude Opus 5.5+12US$ 5.85US$ 0.49
Kimi K3 → Claude Fable 5.1+9US$ 5.63US$ 0.63

This is how any frontier behaves: the last points are the most expensive. It does not mean they are not worth it. It means they are worth it only on the task where those points make a difference.

What does cost per point not show?

Speed, task type and effort. The math above is a snapshot, not a verdict.

Blind spotWhy it matters
SpeedMiMo-V2.6-Pro generates 41 tokens per second, versus 87 for GLM-5.3 and 93 for Opus 5.5 (Artificial Analysis). Cheap and slow adds up in a long agentic session.
General index, not just codeThe Intelligence Index aggregates 10 evaluations. A model can do better at code than its overall score suggests, and vice versa.
Maximum effortEvery model was measured at the highest effort. At lower effort, token usage drops and the math changes.
Lab pricingCost uses the official API price. Another provider serving the same open model may charge differently.

How do you run the math with your own numbers?

Swap the values for what you measure in your own usage: the score on the benchmark that matters for your work and the real cost per task. The script below runs on Python 3 with no dependencies.

# Cost per point: how much each score point costs, and how much the extra point costs
# Data: Artificial Analysis Intelligence Index, maximum effort, 2026-09-29
models = {
    # name: (index score, average cost per task in US$)
    "MiMo-V2.6-Pro": (46, 0.13),
    "GLM-5.3": (45, 2.01),
    "Kimi K3": (44, 2.00),
    "Claude Opus 5.5": (58, 5.98),
    "Claude Sonnet 5.5": (56, 7.60),
    "Claude Fable 5.1": (53, 7.63),
}

for name, (score, cost) in models.items():
    print(f"{name}: US$ {cost / score:.4f} per point")

# Marginal point: how much you pay for each point gained by moving up a model
base, target = "GLM-5.3", "Claude Opus 5.5"
(sb, cb), (st, ct) = models[base], models[target]
print(f"{base} -> {target}: US$ {(ct - cb) / (st - sb):.2f} per extra point")

To find the cost per task of your own usage, add input and output for a typical task:

def cost_per_task(input_tokens, output_tokens, input_price, output_price):
    # prices in US$ per million tokens
    return (input_tokens * input_price + output_tokens * output_price) / 1_000_000

So which should you pick: open or closed model?

Both, at different moments. The open model for volume, the closed one for the task where the extra points solve the problem.

SituationPickWhy
Common bug, small feature, tests, local refactorOpen (GLM-5.3, Kimi K3)2.3x lower cost per point and a score good enough for the task
Large batch, no rushCheapest open model (MiMo-V2.6-Pro)US$ 0.13 per task, slowness does not hurt
Bug that already failed twice, architecture with many dependenciesClosed (Claude Opus 5.5)This is where the 12 to 13 points show up, and US$ 0.31 per extra point pays for itself
Fixed choice of Sonnet 5.5 or Fable 5.1 at maximum effortRevisitOn the index, Opus 5.5 scores more points for less money per task

In practice, this means switching models mid-session. In Verboo Code, from the terminal:

/model          # opens the list of models available on your plan
/model <id>     # switches directly, the session history carries over
/effort max     # raises effort only for the hard task
/effort auto    # goes back to the model default

The whole cost-per-point exercise exists because every token is billed, and tokens are what make a verbose model cost more than the price list promises. On Verboo Code plans run with unlimited tokens, so raising effort or letting the model think longer changes the response speed, not your bill.

Read also

Enjoyed this article?
Share knowledge with your network.
// Read also

Related articles