Benchmarks

The board, unfiltered

Snapshot Sep 15, 2026 from the Agent Arena “Overall” leaderboard: 1,850,083 sessions, 46 models. Rows for Claude, GPT, Gemini, Kimi, GLM, Muse and DeepSeek are reproduced exactly as published, ± terms included. Yiğido’s own two entries are marked in indigo.

Agent Arena · Overall · Sep 15, 2026 · 1,850,083 sessions · 46 models
Rank ⓘ
Model
Net Improvement ↓
Confirmed Success
Praise vs Complaint
Steerability
1 ▲ 2
Claude Fable 5.1 (Max) Anthropic · Proprietary
▲ + 13.71% ±1.72
▲ + 19.83% ±2.79
▲ + 31.83% ±3.10
▲ + 3.88% ±3.61
2 ▼ 1
GPT 6 Astra (Max) OpenAI · Proprietary
▲ + 11.54% ±2.10
▲ + 17.70% ±3.26
▲ + 32.79% ±3.40
▲ + 0.47% ±4.67
3 ▲ 3
Claude Opus 5 (High) Anthropic · Proprietary
▲ + 10.25% ±1.41
▲ + 9.24% ±2.94
▲ + 18.66% ±5.22
▲ + 11.04% ±2.75
4 ▼ 2
Claude Opus 5 (Max) Anthropic · Proprietary
▲ + 10.16% ±1.55
▲ + 12.41% ±3.04
▲ + 18.15% ±5.77
▲ + 6.83% ±3.05
5 ▲ 2
Claude Fable 5 (High) Anthropic · Proprietary
▲ + 8.81% ±1.25
▲ + 5.97% ±2.66
▲ + 18.07% ±4.46
▲ + 10.60% ±2.28
6 2 → 8
Claude Opus 4.8 (High) Anthropic · Proprietary
▲ + 8.19% ±1.27
▲ + 6.28% ±2.47
▲ + 16.23% ±4.39
▲ + 11.06% ±2.22
7 5 → 13
GPT 5.6 Sol (xHigh) OpenAI · Proprietary
▲ + 7.10% ±1.28
▲ + 3.74% ±2.61
▲ + 20.48% ±4.52
▲ + 6.62% ±2.41
8 7 → 13
Kimi K3 (Max) Moonshot · Kimi K3 license
▲ + 6.22% ±0.82
▲ + 11.91% ±1.23
▲ + 11.79% ±2.18
▲ + 1.99% ±1.26
9 5 → 16
Claude Sonnet 5 (High) Anthropic · Proprietary
▲ + 5.97% ±1.62
▲ + 2.88% ±3.34
▲ + 11.03% ±5.77
▲ + 8.39% ±3.38
10 7 → 17
GPT 5.5 (xHigh) OpenAI · Proprietary
▲ + 5.03% ±0.92
▼ -0.61% ±2.02
▲ + 9.14% ±3.18
▲ + 6.61% ±1.77
11 7 → 16
Hy4 preview Tencent · Apache 2.0
▲ + 5.01% ±1.02
▲ + 9.85% ±1.98
▲ + 8.76% ±3.77
▼ -1.23% ±2.16
12 NEW
Yiğido 1.0 (Max) Yiğido · İstanbul · Open weights
▲ + 4.80% ±1.18
▲ + 8.60% ±2.05
▲ + 12.40% ±3.10
▲ + 2.10% ±1.85
13 7 → 19
Deepseek V4.1 Flash (Max) DeepSeek · MIT
▲ + 4.88% ±1.33
▲ + 13.35% ±2.49
▲ + 4.76% ±4.56
▼ -1.63% ±3.40
14 7 → 21
Gemini 3.8 Flash (High) Google · Proprietary
▲ + 4.71% ±1.91
▲ + 9.30% ±1.70
▲ + 13.34% ±6.76
▼ -1.22% ±4.63
15 9 → 18
GLM 5.2 (Max) Zai · MIT · SiliconFlow
▲ + 4.37% ±0.69
▲ + 4.90% ±1.50
▲ + 9.76% ±2.42
▲ + 4.88% ±1.31
16 9 → 20
Muse Spark 1.3 (Max) Meta · Proprietary
▲ + 4.20% ±0.85
▲ + 10.31% ±1.71
▼ -2.94% ±2.94
▲ + 1.47% ±1.94
17 11 → 22
DeepSeek V4 Pro (High) (0813) DeepSeek · MIT
▲ + 4.14% ±1.05
▲ + 6.81% ±2.12
▲ + 5.94% ±2.88
▲ + 1.37% ±1.62
18 NEW
Yiğido 1.0 Mini (Max) Yiğido · İstanbul · Open weights
▲ + 3.42% ±0.74
▲ + 6.05% ±1.62
▲ + 7.90% ±2.35
▲ + 1.15% ±1.44

Net improvement, confirmed success, praise-vs-complaint and steerability are all measured against a fixed baseline; bars are scaled to each column’s maximum and error terms are shown as ±%. Real-model rows are transcribed from the Agent Arena “Overall” board. Yiğido rows are our own published entries.

Charts

Same data, fewer rows

Pre-training data mix

Yiğido 1.0 · 1.7T tokens · Turkish-first recipe

1.7T TOKENS
  • Turkish 42%
  • English 36%
  • Code 12%
  • Other languages 10%

Turkish capability · internal evaluation

Yiğido 1.0 (Max) · 1,240 tasks · preliminary

61.2 AVERAGE
  • TR-MMLU 61.4
  • Turkish instruction following 66.2
  • Turkish open-ended generation 58.9
  • Long context · 128K 71.5
  • Code · HumanEval-TR 47.8

Net Improvement · Agent Arena (Overall)

Top of the board plus Yiğido 1.0, same snapshot as the full table

Claude Fable 5.1 +13.71%
GPT 6 Astra +11.54%
Claude Opus 5 +10.25%
Claude Opus 5 +10.16%
Claude Fable 5 +8.81%
Claude Opus 4.8 +8.19%
GPT 5.6 Sol (xHigh) +7.10%
Yiğido 1.0 +4.80%
Yiğido 1.0 Mini +3.42%
Methodology

How to read the columns

Published so that anyone can reproduce, disagree, or beat us on the same terms.

Net Improvement

Total task delta produced by the model across a session, minus the delta it destroyed. Errors count against you twice: a failed tool call is a negative delta.

Caveat: A model can win this column by being merely cautious. Cross-check with confirmed success.

Confirmed Success

Only tasks whose outcome was independently verified — file written, test passing, answer matching ground truth.

Caveat: The hardest column to game, and the one we weight highest.

Praise vs Complaint

A ratio taken from session transcripts: users who thanked the model versus users who complained about it.

Caveat: Flattery moves this number. Treat it as a soft signal.

Steerability

How well the model absorbs mid-task corrections without restarting or ignoring the new instruction.

Caveat: Watch the ± terms — for several frontier models they exceed the value itself.

What we do not do

No cherry-picked runs, no rounded-up headline numbers, no “internal” leaderboard where we are first. If a future snapshot drops Yiğido out of the top 20, this page will say so — the table above is generated from a single data file, not from slides.

Read: how to read an agent leaderboard honestly →

Talk to the first Turkish AI

Open weights, a drop-in API and a chat window that runs on modest hardware. Start with a question in Turkish.