How to read an agent leaderboard honestly
Error terms, session counts and why a 4.8% net improvement is not a rounding accident.
The first Turkish foundation model to ship weights, a datasheet and a leaderboard entry at the same time.
Yiğido 1.0 is live. Two models, one recipe, everything published:
Weights, tokenizer, data datasheet and the evaluation harness are in the release. No research-only clause, no contact form. If you run a bank in Istanbul with strict data-residency rules, you can host Yiğido inside your own network.
On the Agent Arena “Overall” snapshot of 15 September 2026, Yiğido 1.0 (Max) enters at rank 12 with +4.80% net improvement, 8.60% confirmed success and 2.10% steerability. That places it above several closed frontier models on confirmed success, and behind the top Anthropic and OpenAI runs.
Rank 12 is the honest number, and it is the number we will keep publishing — including the ± error terms, including the runs where we lose.
The 1.1 branch targets three things: longer Turkish documents, better tool-call reliability, and a cheaper Mini that still fits your laptop.
Error terms, session counts and why a 4.8% net improvement is not a rounding accident.
Agglutination, vowel harmony and the quiet tax that general-purpose tokenizers put on every Turkish sentence.