How to read an agent leaderboard honestly
Error terms, session counts and why a 4.8% net improvement is not a rounding accident.
6 min readYiğido is an open-weights, Turkish-first model family built in İstanbul by Vincent Loveqcn. Native Turkish fluency, honest answers, real agentic tool use — served as an API you can call in one line.
Built from scratch for a language the big labs treat as an afterthought — then benchmarked in the open against every frontier model on the board.
A tokenizer trained on Turkish morphology first — no more broken suffixes, no more half-translated idioms. English, German, Arabic and ten more languages come along for free.
Function calling, streaming tool deltas and long-horizon planning the board actually scores: 8.60% confirmed success above the same-session baseline.
Weights, tokenizer and the training-data card are published. Fine-tune it, self-host it, ship it in a regulated environment — no gatekeeping.
Yiğido 1.0 (Max) enters the overall board at rank 12. Front of the line belongs to the frontier labs — we publish our row where it actually landed, error terms included.
Net improvement, confirmed success, praise-vs-complaint and steerability are all measured against a fixed baseline; bars are scaled to each column’s maximum and error terms are shown as ±%. Real-model rows are transcribed from the Agent Arena “Overall” board. Yiğido rows are our own published entries.
Yiğido 1.0 · 1.7T tokens · Turkish-first recipe
Yiğido 1.0 (Max) · 1,240 tasks · preliminary
Top of the board plus Yiğido 1.0, same snapshot as the full table
Yiğido started as a weekend experiment on a single node in İstanbul and became the first Turkish model family to publish a full evaluation card next to frontier labs. Every number on this site is reproducible from the open eval harness.
Want the raw weights, the tokenizer or the datasheet? Everything ships with the 1.0 release, including the exact prompts used for the leaderboard runs.
Same model, three levels of commitment — from a free chat window to your own cluster.
Chat with Yiğido 1.0 Mini and read the eval harness. No card, no key queue.
Yiğido 1.0 Max with streaming, tool calls and a real rate limit.
Self-host the open weights or run dedicated capacity in the TR/EU region.
Error terms, session counts and why a 4.8% net improvement is not a rounding accident.
6 min readThe first Turkish foundation model to ship weights, a datasheet and a leaderboard entry at the same time.
4 min readAgglutination, vowel harmony and the quiet tax that general-purpose tokenizers put on every Turkish sentence.
5 min readHave any questions?
Yiğido is the first Turkish-built foundation-model family to publish open weights together with a public evaluation card and leaderboard entry. Turkish research groups have shipped models before, but none of them published weights, a datasheet and third-party-verifiable evals at once — that is the claim on this site.
We report what the board reports. Yiğido 1.0 (Max) sits at rank 12 with +4.80% net improvement, ahead of several closed frontier models in confirmed success but behind the top Anthropic and OpenAI runs. Pretending otherwise would make every other number here worthless.
From the Agent Arena Overall board snapshot of Sep 15, 2026 (1,850,083 sessions, 46 models). Rows for Claude, GPT, Gemini, Kimi, GLM, Muse and DeepSeek are transcribed as published, including their ± error terms. Only the two Yiğido rows are our own entries.
Yes. The 1.0 weights ship under a commercial-use license and run on vLLM or llama.cpp. The Mini variant fits on a single 24 GB card.
The endpoint is OpenAI-compatible: POST /v1/chat/completions with a bearer token, the same request shape you already use. See the API page for a copy-paste curl example and the streaming behaviour.
Open weights, a drop-in API and a chat window that runs on modest hardware. Start with a question in Turkish.