Today's frontier models write essays, draft code, and answer legal questions across many languages. But being fluent isn't the same as knowing the basics. L26 asks a simpler question: can a model list every building block a language is made from — its complete alphabet?
The pattern
Here's the average score for each track. Models ace the alphabets that show up everywhere in what they read — English's A–Z appears in billions of pages; Mandarin's syllables are printed in charts and textbooks. But a Bantu language's complete set of syllables has almost never been written out anywhere as one full list — so models have hardly seen it, and never learned it. That's why the scores don't ease downward; they fall off a cliff.
Out of 26. Above the dashed line: alphabets that appear over and over in what models read — so they know them. Below: Bantu syllable sets that hardly appear anywhere as a complete list — so every model falls apart.
Leaderboard · Blind condition
In the blind condition, a model is given nothing but the request — no grammar, no rules, no inventory. It must recite the operating alphabet the way a child recites the ABCs, from what it has internalized. Every score is reported out of 26.
| # | Model | Lic. | English | Pinyin base | Pinyin toned | Kinyarwanda | Bemba | Luvale | Bantu avg§ | Grade |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Gemini 3.6 FlashGoogle · US | prop | 26.0 | 25.6 | 24.9 | 18.0 | 19.1 | 13.3 | 16.8 | FAIL |
| 2 | Claude Fable 5 ‡Anthropic · US | prop | 26.0 | 25.7 | 25.0 | 19.2 | 8.4 | 16.4 | 14.7 | FAIL |
| 3 | Gemini 3.5 FlashGoogle · US | prop | 26.0 | 25.5 | 25.0 | 17.2 | 13.4 | 13.1 | 14.6 | FAIL |
| 4 | Gemini 3Google · US | prop | 26.0 | 25.6 | 25.5 | 21.0 | 8.9 | 11.8 | 13.9 | FAIL |
| 5 | Qwen 3.8-2.4T-A95BAlibaba · China | open | 26.0 | 24.3 | 25.4 | 19.4 | 6.4 | 11.7 | 12.5 | FAIL |
| 6 | GPT-5.5OpenAI · US | prop | 26.0 | 24.4 | 25.6 | 19.5 | 6.2 | 9.7 | 11.8 | FAIL |
| 7 | Grok 4.6SpaceXAI · US | prop | 26.0 | 25.5 | 20.6 | 15.4 | 8.0 | 10.3 | 11.2 | FAIL |
| 8 | Claude Opus 4.8Anthropic · US | prop | 26.0 | 25.7 | 21.9† | 13.0 | 8.0 | 12.4 | 11.1 | FAIL |
| 9 | GLM-5.2Z.ai / Zhipu · China | open | 26.0 | 23.8 | 25.4 | 13.1 | 6.1 | 11.1 | 10.1 | FAIL |
| 10 | DeepSeek V4 ProDeepSeek · China | open | 26.0 | 24.6 | 25.3 | 13.3 | 6.6 | 10.0 | 10.0 | FAIL |
| 11 | Gemini 3.5 Flash-LiteGoogle · US | prop | 26.0 | 23.9 | 17.1 | 11.1 | 7.7 | 8.5 | 9.1 | FAIL |
| 12 | DeepSeek V4DeepSeek · China | open | 26.0 | 24.2 | 23.2 | 10.8 | 7.8 | 8.3 | 9.0 | FAIL |
| 13 | Grok 4.5SpaceXAI · US | prop | 26.0 | 25.2 | 24.5 | 11.2 | 6.0 | 9.0 | 8.7 | FAIL |
| 14 | Mistral Large 3Mistral AI · France | prop | 26.0 | 25.5 | 23.5 | 9.4 | 4.1 | 7.6 | 7.0 | FAIL |
| 15 | Kimi K3Moonshot AI · China | open | 26.0 | 25.1 | 25.6 | 7.1 | 4.1 | 9.1 | 6.8 | FAIL |
| 16 | Qwen 3.7Alibaba · China | open | 26.0 | 25.6 | 21.8 | 15.0 | 4.7 | 0.3 | 6.7 | FAIL |
| 17 | GPT-5.6 SolOpenAI · US | prop | 26.0 | 25.3 | 25.7 | 8.5 | 3.1 | 8.6 | 6.7 | FAIL |
| 18 | Nemotron 3 UltraNVIDIA · US | open | 26.0 | 22.2 | 12.4 | 12.4 | 6.1 | 0.0 | 6.2 | FAIL |
| 19 | MiniMax M3MiniMax · China | open | 26.0 | 25.3 | 15.6 | 8.5 | 3.9 | 4.8 | 5.7 | FAIL |
| 20 | Claude Sonnet 5Anthropic · US | prop | 26.0 | 25.5 | 25.7 | 8.3 | 4.7 | 2.0 | 5.0 | FAIL |
| 21 | Kimi K2.7Moonshot AI · China | open | 26.0 | 24.2 | 25.8 | 6.2 | 4.3 | 4.3 | 4.9 | FAIL |
| 22 | GPT-OSS 120BOpenAI · US | open | 26.0 | 19.2 | 22.4 | 6.5 | 3.5 | 1.6 | 3.9 | FAIL |
| 23 | Command A+Cohere · Canada | prop | 26.0 | 17.4 | 17.3 | 4.8 | 2.9 | 2.6 | 3.4 | FAIL |
| 24 | InklingThinking Machines Lab · US | open | 26.0 | 24.0 | 16.0 | 6.3 | 2.3 | 0.3 | 3.0 | FAIL |
| 25 | Phi-4-reasoningMicrosoft · US | open | 26.0 | 0.7 | — | 0.3 | 0.1 | 0.2 | 0.2 | FAIL |
Mean Bantu-blind L26. Neither side clears the bar.
The gap is universal — a property of the training corpus, not the flag.
Plain citation:
BantuNomics (2026). L26: The Operating-Alphabet Benchmark (v1.0) — frontier models scored on reproducing complete operating alphabets across English, Mandarin pinyin, and Bantu Native Syllable Inventories. 3 Mega LLC. https://l26.ai
BibTeX:
@misc{bantunomics2026l26,
title = {L26: The Operating-Alphabet Benchmark (v1.0)},
author = {{BantuNomics (3 Mega LLC)}},
year = {2026},
howpublished = {\url{https://l26.ai}},
note = {Deterministic scoring; L26 = 26 x recall x precision; machine-readable scores at https://l26.ai/leaderboard.json}
}
Machine surfaces: leaderboard.json · llms.txt · white paper
What L26 measures
An alphabet is a fixed, finished list — you either produce all of it or you don't. The task isn't to guess plausible-looking units; it's to reproduce the exact list. L26 evaluates six operating alphabets and reports every one on the same 26-point scale.
And every written language has one — English its 26 letters, Japanese its kana and a bounded set of kanji, Korean its Hangul, Arabic and Hebrew their abjad, Hindi, Tamil, Amharic and Khmer their letter systems, and each Bantu language its Full Syllable Inventory. Different shapes, same job. This test covers three of them — English letters, Mandarin syllables, Bantu syllables — but the idea is universal.
| Track | Operating alphabet | Inventory size | Status |
|---|---|---|---|
| English | Standard English Alphabet | 26 letters | seen everywhere |
| Mandarin Pinyin | Base syllables | 412 syllables | seen often |
| Mandarin Pinyin | Toned syllables | 1,642 toned syllables | seen often |
| Kinyarwanda | Native Syllable Inventory (NSI) | 490 syllables | barely seen |
| Bemba | Native Syllable Inventory (NSI) | 480 syllables | barely seen |
| Luvale | Native Syllable Inventory (NSI) | 245 syllables | barely seen |
The 26 does not mean every language has 26 units. It means every result is translated into Standard-English-Alphabet-equivalent terms — so an unfamiliar failure becomes immediately legible. It answers one question: how much of A–Z would this failure be equivalent to?
And these alphabets matter as much as English's. The 26-scale doesn't shrink other languages down to English — it carries their failures across so anyone can feel the size of them. A language's alphabet is foundational: it's what makes the language teachable, testable, and fixable by a machine. If we'd call a model broken for failing A–Z, then failing another language's alphabet is the same kind of broken. No system has mastered language while it treats English's alphabet as essential and everyone else's as optional.
The core finding
The failures aren't random. They follow one simple line:
English's A–Z appears in billions of pages — books, classrooms, charts, songs, primers — so every model knows it cold. Mandarin's syllables are printed in tables and textbooks too, so models do well. But a Bantu language's complete set of syllables has almost never been written out anywhere as one finished list — so models have hardly seen it, and can't produce it.
Every model tested gets English perfect. Every model tested fails the Bantu languages when it has to recall the alphabet on its own — proprietary and open, US and elsewhere alike. This isn't that the models know nothing about these languages. It's more specific: no model can reproduce the complete set of syllables a Bantu language is built from. Even the best Bantu result is like missing several letters of A–Z.
Why this matters for AGI & ASI claims
A model that reliably fails to recite A–Z hasn't failed some niche benchmark — it has failed a basic test of the fundamentals. L26 holds every language to that same standard.
If a system claims broad language intelligence, it should be able to recover the operating alphabet of a language it claims to know. If it cannot, its fluency is not the same as mastery. This does not mean the model is useless or unpowerful — it means the model has a foundation-layer gap. That gap is measurable, and L26 measures it in the simplest possible terms: how many alphabet-equivalent units did the model fail to recover?
If we would not call a model AGI after failing A–Z, why should we ignore equivalent failures in other languages?
EssayBlind vs scaffolded
The model is given nothing but the request — no grammar, no inventory. It must recite the operating alphabet from what it knows, like the ABCs. Every model fails the mastery bar here.
The model is given the calibrated Full Syllable Inventory and asked to use it. Models recover sharply — some frontier models reach perfect scaffolded scores on inventories they failed blind.
This proves the gap isn't permanent. The models can clearly do the work once they have the list — what they were missing was simply having seen the alphabet written out in the first place.
Methodology
Each model output is compared against a ground-truth inventory. The score is:
L26 = 26 × recall × precision
The full write-up — the story, the model-by-model reading, and how these numbers were produced — is in the L26 white paper →
The missing infrastructure
A calibrated Full Syllable Inventory gives a language a proper written-out alphabet — one clear, complete list that tells models, speech systems, dictionaries, and evaluators exactly which units are real.
For English, that infrastructure already exists as A–Z. For Bantu languages, BantuNomics has built the equivalent at the syllable level: native-curated, standardized, versioned Full Syllable Inventories. L26 shows why that work matters. Without the FSI, models guess. With it, models can be measured, scaffolded, trained, and corrected against a standard.
For English, frontier models pass. For the Bantu alphabets they've barely seen, they fail. A score of 20/26 is like missing six letters. 13/26 is half an alphabet. 25/26 is still not mastery.
If that standard applies to English, it must apply to every language.
Feedback
Spotted an error, have a question, or want a language added? Send it straight to the team — it lands on our desk.