BenchmarksMethodology

How to read an agent leaderboard honestly

Error terms, session counts and why a 4.8% net improvement is not a rounding accident.

18 September 2026 · 6 min read · Vincent Loveqcn · Yiğido Lab, İstanbul

Every number on our benchmark page is copied from the Agent Arena “Overall” snapshot of 15 September 2026 — 1,850,083 sessions across 46 models. Four columns matter:

  • Net improvement — the deltas a model produces on real agent tasks, minus the deltas that made things worse.
  • Confirmed success — tasks where the outcome was independently checked, not self-reported.
  • Praise vs complaint — a sentiment ratio taken from the session transcripts.
  • Steerability — how well the model follows mid-task corrections.

Read the ± as carefully as the number

GPT 6 Astra scores 0.47% steerability ± 4.67. Claude Fable 5.1 scores 3.88% ± 3.61. Two frontier models, and both error terms are large enough to make the ordering of neighbouring rows almost meaningless in any single snapshot. That is not a flaw in the board; it is the truth about measuring agents.

The one place a model can fake progress

Praise-vs-complaint rewards models that compliment the user. Net improvement is harder to fake because a broken tool call is a broken tool call. When you compare models, weight the middle two columns and treat the far-left column as the headline it is.

We will keep publishing our own row next to the ones we lose on, with error bars attached.