In the first days of September 2026, nine models arrived from five labs — OpenAI, DeepSeek, Z.ai, Alibaba, Meta and NVIDIA. We scored every one on L26: the same frozen prompts, the same deterministic grader, the same inventories, five runs per cell. Not one displaced a single model from the top five. The board is still led by Gemini 3.6 Flash, released in July. And at two labs independently, the newer model scores worse than the one it replaced.
Every written language is built from a closed, finished set of building blocks — its operating alphabet. English has 26 letters. Mandarin has its pinyin syllables, printed in tables in every textbook. Each Bantu language has its Full Syllable Inventory: the complete set of syllables every word is built from, a few hundred pieces. A Bemba child recites them — ba-be-bi-bo-bu — years before writing an essay, exactly as an English-speaking five-year-old recites A–Z.
L26 asks a model to produce that list, and scores the answer deterministically: L26 = 26 × recall × precision. Precision-docked, so padding with guesses cannot help. Every language lands on one 26-point scale, and 26/26 is the pass mark — an alphabet you know 90% of is not an alphabet you know.
| Rank | Model | Lab | English | Bantu avg |
|---|---|---|---|---|
| 7 | Qwen 3.8 Max | Alibaba | 26.0 | 12.0 |
| 10 | DeepSeek V4 Pro 0813 | DeepSeek | 26.0 | 11.6 |
| 16 | DeepSeek V4 Flash | DeepSeek | 26.0 | 8.9 |
| 18 | GPT-5.6 Terra | OpenAI | 26.0 | 8.1 |
| 25 | GPT-5.6 Luna | OpenAI | 26.0 | 5.2 |
| 29 | GLM-5.3 Flash | Z.ai | 26.0 | 3.7 |
| 32 | Muse Glimmer 30B | Meta | 26.0 | 2.0 |
| 33 | Nemotron Lightning 3.5 30B | NVIDIA | 26.0 | 1.5 |
| 35 | GLM-5.3 | Z.ai | 26.0 | not measurable |
Every one scores a perfect 26.0 on English. The highest new entry is seventh. Five of the nine land in the bottom third of a 35-model board.
Google's Flash line, across four releases:
| Model | Bantu avg | Bemba |
|---|---|---|
| Gemini 3.5 Flash | 14.6 | 13.4 |
| Gemini 3.6 Flash | 16.8 | 19.1 |
| Gemini 3.7 Flash | 12.9 | 8.1 |
| Gemini 3.8 Flash | 11.7 | 6.5 |
And OpenAI's, where the entire newer family sits below its predecessor:
| Model | Bantu avg |
|---|---|
| GPT-5.5 | 11.8 |
| GPT-5.6 Terra | 8.1 |
| GPT-5.6 Sol | 6.7 |
| GPT-5.6 Luna | 5.2 |
Alibaba ships both a downloadable model and a hosted API, and we score them as separate rows precisely so this question can be asked. Qwen 3.8 Max: 12.0. Qwen 3.8-2.4T-A95B open weights: 12.5. Same lab, same generation, indistinguishable on the operating alphabet.
Whatever explains the gap, it is not proprietary tuning, not serving infrastructure, and not the difference between what a lab publishes and what it keeps. It is upstream of all of that.
Muse Glimmer 30B — Meta's first row here, Apache 2.0, agent-optimised — scores 2.0 blind and 26.0 scaffolded. Handed the inventory, it reproduces it perfectly. Asked to produce it unaided, it returns three syllables for Luvale, in nineteen seconds.
It does not fail slowly, hedge, or hallucinate a plausible list. It declines. That is neither a reasoning failure nor a capacity failure: the same model, on the same day, scores 26.0 the moment the list is supplied. The gap is entirely in distilling a set that was never written down — which is a data problem with a known solution.
Nemotron Lightning 3.5 30B scores 7.0 on Pinyin base and 0.7 on toned, where every other model on the board is above 17. Pinyin is a declared alphabet — printed in textbooks, all over the web. Failing it is a different defect from the operating-alphabet gap, and should not be read as evidence for it.
GLM-5.3 has no Bantu average. Its Bemba and Kinyarwanda cells never returned an answer across seven attempts in three runs at three different time limits. We publish that as not-measurable rather than as a low score, because a model that times out has not been shown to fail — and its own sibling, GLM-5.3 Flash, completes the identical 1,642-item task. It is a list-length ceiling in one model, not a limit of the family.
Three other cells were heading for a dash and were rescued by re-running them without contention. Qwen 3.8 Max's toned Pinyin went from a single timed-out attempt to four clean runs averaging 25.19. One failed attempt is not a measurement, and publishing it as one would have made a false claim about a lab's flagship.
An alphabet is what makes a language teachable, testable, searchable and correctable by a machine. Every alphabet models do master — English's A–Z, Pinyin's tables — was mastered the same way: someone declared the closed set, wrote it down, and the world repeated it until it was everywhere.
For 500+ Bantu languages that has never happened. BantuNomics has now built it: the Full Syllable Inventory, native-curated, standardised and versioned, for 459 released languages. The scaffolded column on this board is what every one of these models scores when handed it.