AAI Labs
Main
Services
CasesResearchTeam
Company
More
Contact
Loading…

Services

  • AI for Energy
  • AI for Municipalities
  • AI for Transport
  • Generative AI
  • LLM for Business
View all services →

Company

  • About
  • Careers
  • Cases
  • Privacy Policy
  • Public R&D
  • Contact

More

  • EU AI Act hub
  • AI dictionary
  • Research
  • Blog
  • News

Products

  • AI Team for Hire
  • Merkys.AI
  • Cargobroker.AI
  • Klikt

UAB Taikomasis dirbtinis intelektas © 2026

[email protected]LinkedIn →
Blog

Which open-source LLM is cheapest to trust for AI agents?

October 2, 2026
6 min read

In July 2026, the Berkeley Function Calling Leaderboard (BFCL), one of the most widely cited benchmarks for AI tool use, ranked Qwen3 235B as the strongest open-weight model for calling tools, at about 88.5% accuracy. It's a useful starting point. However, it scores each call on its own and says nothing about what a model costs to run, whether it stays consistent over a long task, or what hardware it needs.

Those gaps matter most when you build AI agents on open-weight models, the ones you can host yourself. Agents rarely fail because they write a weak paragraph; usually, failures are caused by calling the wrong tool, inventing a parameter or getting the context mixed up halfway through a fifteen-step task. Every failure requires a retry or a person stepping in, and both cost money.

The difference between being right once and being right every time is what makes a model reliable. In the original τ-bench study, GPT-4o completed about 61% of retail customer-service tasks on a single attempt. When it had to succeed in all eight repeated runs of the same task, that fell to under 25%.

So during our July mini-projects, one team asked a practical question: for a company running its own agents, which open-weight model is the cheapest one to trust?

Why price per token is the wrong number

The team's headline measure is cost per successful task: what one task costs in tokens, divided by how often the model completes it. A model that fails has to try again, so its real cost per finished task is higher than its price list suggests. If one attempt costs $0.03 and the model succeeds 60% of the time, each successful task effectively costs $0.05.

To make the models comparable, every cost is calculated for the same standard agent task: about 15 tool calls, using around 60,000 input tokens and 6,000 output tokens.

The shift from tracking tokens to tracking cost per outcome shows what a finished task really costs.

How the 13 models were chosen

The team started from compar:IA, the French government's public AI comparison service, which listed 113 models in its June 2026 update. They used it as a catalogue only, not as a ranking, and applied four filters:

  • Open weights only, so the model can be self-hosted.
  • One current flagship per model family, not every point release.
  • Native tool-calling support.
  • At least some published agent benchmark results.

That left 13 models. Each was scored from public sources: tool-calling accuracy from the Berkeley Function Calling Leaderboard (BFCL); reliability from BFCL's multi-turn tests, its test of whether a model knows when not to call a tool, and τ²-bench; prices from provider pricing pages. Every number carries its source and retrieval date.

One practical note: some of these models are released under "semi-open" licenses, and their commercial terms differ. Check the license before you deploy.

What an agent task actually costs (July 2026 data)

For the same standard task, the cheapest model costs $0.008 and the most expensive $0.127, a 16x difference. These are the eight models with full cost data, ranked by cost per successful task:

#ModelCost per successful taskReliabilityHardware needed
1Qwen3 32B$0.00881.2%Single GPU (80 GB)
2MiniMax M3$0.02888.9%Multi-node cluster
3DeepSeek V4 Pro$0.03980.7%Multi-node cluster
4GLM-4.5-Air$0.04046.5%Single GPU (80 GB)
5Qwen3 235B$0.04683.9%Single server (8×80 GB)
6GLM-5 / 5.1$0.05389.7%Multi-node cluster
7Kimi K2 / K2.6$0.05780.9%Multi-node cluster
8Mistral Large 3$0.12730.7%Single server (8×80 GB)

Three things stand out.

The cheapest model is also the best value. Qwen3 32B is a small model that runs on a single 80 GB GPU. It costs less than a cent per task and still reaches 81.2% reliability.

High reliability doesn't have to be expensive. MiniMax M3 reaches 88.9% reliability, close to the top of the list, at $0.028 per task.

The most expensive model is the least reliable. Mistral Large 3 costs $0.127 per task with 30.7% reliability, the lowest of the eight.

Getting one call right is not the same as being reliable

Tool-calling accuracy measures whether a model picks the right function and fills in valid arguments. Reliability measures whether it keeps doing that across a multi-step task, knows when not to call a tool, and succeeds consistently across repeated runs.

GLM-4.5-Air scores 76.7% on tool calling but only 46.5% on reliability. GLM-5 is the reverse: 62.1% on tool calling, but the most reliable model in the comparison at 89.7%.

Choose by tool-calling scores alone and you might pick GLM-4.5-Air, then watch your agent break halfway through a task. Choose by reliability alone and you might pass over a model that handles short, single-step work well. That's why the comparison weighs both.

GLM-4.5-Air vs GLM-5: higher tool-calling accuracy does not mean higher reliability

No model wins everything

CategoryWinnerResult
Cheapest per task and best valueQwen3 32B$0.008 per task
Most reliableGLM-5 / 5.189.7%
Best tool callingQwen3 235B88.5%
FastestLlama Nemotron 70B277 tokens per second

The fastest model isn't the cheapest, and the most accurate isn't the most reliable. Decide which of these matters most for your agent before you look at the ranking.

Start with the hardware you have

For most companies, the first question isn't which model is best. It's which one they can actually run.

  • One GPU (80 GB): Qwen3 32B is the clear choice. GLM-4.5-Air runs on the same hardware but costs five times as much per task and is far less reliable.
  • One server (8 × 80 GB GPUs): Qwen3 235B costs $0.046 per task, reaches 83.9% reliability and has the best tool-calling score in the comparison. Mistral Large 3 needs similar hardware, costs almost three times as much and is much less reliable.
  • A multi-GPU cluster or an API: MiniMax M3 is the cheapest high-reliability option. If reliability matters more than cost, GLM-5 is the most reliable model, at about twice the price.
Open-weight models compared by tool calling, reliability, speed, cost per task and hardware needed

Longer tasks change the numbers. Cost grows with every token, and every extra step is another chance for an unreliable model to fail. For long-running agents, reliability should count for more than price.

What these numbers can and can't tell you

  • The scores come from public benchmarks and provider prices, not from our own tests. The team's contribution is putting scattered sources side by side and adding cost and hardware.
  • Public benchmark scores are a best case. Models are increasingly tuned for well-known benchmarks, so expect weaker results on your own tasks.
  • Not every model is covered by every benchmark. Some reliability scores rest on one source, others on three.
  • Models and prices change every month. This is a snapshot from 10 July 2026.

How to choose a model for your agent

  • Estimate your task size. Count roughly how many tool calls and tokens one task takes.
  • Compare cost per successful task, not price per token.
  • Filter by what you can use: the hardware you have and the license terms you can accept.
  • Test the shortlist on your own tasks before you commit. Public scores narrow the field; your own workflows decide.

A model that is cheap per token but fails often is not a cheap agent. In this comparison, the cheapest model to trust was a small one that runs on a single GPU, not one of the large frontier models. Public benchmarks are a good way to build a shortlist, but the final choice should come from testing on your own tasks.

ON THIS PAGE

  • Which open-source LLM is cheapest to trust for AI agents?
  • Why price per token is the wrong number
  • How the 13 models were chosen
  • What an agent task actually costs (July 2026 data)
  • Getting one call right is not the same as being reliable
  • No model wins everything
  • Start with the hardware you have
  • What these numbers can and can't tell you
  • How to choose a model for your agent

Related articles

Eight hours a month for an MVP: how AI mini-projects bring our teams together

Oct 2, 2026

AI Agents in Business: How Working with Software Will Change

Sep 28, 2026

Reminders Cut No-Shows. So Why Didn't the Waiting List Get Shorter?

Sep 28, 2026