Yiğido 1.0 is out: open weights, published evals
The first Turkish foundation model to ship weights, a datasheet and a leaderboard entry at the same time.
Error terms, session counts and why a 4.8% net improvement is not a rounding accident.
Every number on our benchmark page is copied from the Agent Arena “Overall” snapshot of 15 September 2026 — 1,850,083 sessions across 46 models. Four columns matter:
GPT 6 Astra scores 0.47% steerability ± 4.67. Claude Fable 5.1 scores 3.88% ± 3.61. Two frontier models, and both error terms are large enough to make the ordering of neighbouring rows almost meaningless in any single snapshot. That is not a flaw in the board; it is the truth about measuring agents.
Praise-vs-complaint rewards models that compliment the user. Net improvement is harder to fake because a broken tool call is a broken tool call. When you compare models, weight the middle two columns and treat the far-left column as the headline it is.
We will keep publishing our own row next to the ones we lose on, with error bars attached.
The first Turkish foundation model to ship weights, a datasheet and a leaderboard entry at the same time.
Agglutination, vowel harmony and the quiet tax that general-purpose tokenizers put on every Turkish sentence.