The board, unfiltered
Snapshot Sep 15, 2026 from the Agent Arena “Overall” leaderboard: 1,850,083 sessions, 46 models. Rows for Claude, GPT, Gemini, Kimi, GLM, Muse and DeepSeek are reproduced exactly as published, ± terms included. Yiğido’s own two entries are marked in indigo.
Net improvement, confirmed success, praise-vs-complaint and steerability are all measured against a fixed baseline; bars are scaled to each column’s maximum and error terms are shown as ±%. Real-model rows are transcribed from the Agent Arena “Overall” board. Yiğido rows are our own published entries.
Same data, fewer rows
Pre-training data mix
Yiğido 1.0 · 1.7T tokens · Turkish-first recipe
- Turkish 42%
- English 36%
- Code 12%
- Other languages 10%
Turkish capability · internal evaluation
Yiğido 1.0 (Max) · 1,240 tasks · preliminary
- TR-MMLU 61.4
- Turkish instruction following 66.2
- Turkish open-ended generation 58.9
- Long context · 128K 71.5
- Code · HumanEval-TR 47.8
Net Improvement · Agent Arena (Overall)
Top of the board plus Yiğido 1.0, same snapshot as the full table
How to read the columns
Published so that anyone can reproduce, disagree, or beat us on the same terms.
Net Improvement
Total task delta produced by the model across a session, minus the delta it destroyed. Errors count against you twice: a failed tool call is a negative delta.
Caveat: A model can win this column by being merely cautious. Cross-check with confirmed success.
Confirmed Success
Only tasks whose outcome was independently verified — file written, test passing, answer matching ground truth.
Caveat: The hardest column to game, and the one we weight highest.
Praise vs Complaint
A ratio taken from session transcripts: users who thanked the model versus users who complained about it.
Caveat: Flattery moves this number. Treat it as a soft signal.
Steerability
How well the model absorbs mid-task corrections without restarting or ignoring the new instruction.
Caveat: Watch the ± terms — for several frontier models they exceed the value itself.
What we do not do
No cherry-picked runs, no rounded-up headline numbers, no “internal” leaderboard where we are first. If a future snapshot drops Yiğido out of the top 20, this page will say so — the table above is generated from a single data file, not from slides.
Read: how to read an agent leaderboard honestly →Talk to the first Turkish AI
Open weights, a drop-in API and a chat window that runs on modest hardware. Start with a question in Turkish.