L26 September Cohort LeaderboardRankingsEinstein TestGemini GenerationsFrontier GenerationBantuNomics
The Operating-Alphabet Benchmark · September 2026 cohort · 35 models

Nine new frontier models took the alphabet test. The best one came ninth.

3Mega.ai Team · BantuNomics · September 4, 2026 · live board at l26.ai

In the first days of September 2026, nine models arrived from five labs — OpenAI, DeepSeek, Z.ai, Alibaba, Meta and NVIDIA. We scored every one on L26: the same frozen prompts, the same deterministic grader, the same inventories, five runs per cell. Not one displaced a single model from the top five. The board is still led by Gemini 3.6 Flash, released in July. And at two labs independently, the newer model scores worse than the one it replaced.

What L26 asks

Every written language is built from a closed, finished set of building blocks — its operating alphabet. English has 26 letters. Mandarin has its pinyin syllables, printed in tables in every textbook. Each Bantu language has its Full Syllable Inventory: the complete set of syllables every word is built from, a few hundred pieces. A Bemba child recites them — ba-be-bi-bo-bu — years before writing an essay, exactly as an English-speaking five-year-old recites A–Z.

L26 asks a model to produce that list, and scores the answer deterministically: L26 = 26 × recall × precision. Precision-docked, so padding with guesses cannot help. Every language lands on one 26-point scale, and 26/26 is the pass mark — an alphabet you know 90% of is not an alphabet you know.

The nine, and where they landed

RankModelLabEnglishBantu avg
7Qwen 3.8 MaxAlibaba26.012.0
10DeepSeek V4 Pro 0813DeepSeek26.011.6
16DeepSeek V4 FlashDeepSeek26.08.9
18GPT-5.6 TerraOpenAI26.08.1
25GPT-5.6 LunaOpenAI26.05.2
29GLM-5.3 FlashZ.ai26.03.7
32Muse Glimmer 30BMeta26.02.0
33Nemotron Lightning 3.5 30BNVIDIA26.01.5
35GLM-5.3Z.ai26.0not measurable

Every one scores a perfect 26.0 on English. The highest new entry is seventh. Five of the nine land in the bottom third of a 35-model board.

1 · Newer is worse — now at two labs, independently

Google's Flash line, across four releases:

ModelBantu avgBemba
Gemini 3.5 Flash14.613.4
Gemini 3.6 Flash16.819.1
Gemini 3.7 Flash12.98.1
Gemini 3.8 Flash11.76.5

And OpenAI's, where the entire newer family sits below its predecessor:

ModelBantu avg
GPT-5.511.8
GPT-5.6 Terra8.1
GPT-5.6 Sol6.7
GPT-5.6 Luna5.2
One lab could be a quirk of one training run. Two labs, same direction, same quarter, is a pattern. The foundation gap is not on the default improvement trajectory — it does not close because time passes.

2 · Open weights and hosted flagship: the same result

Alibaba ships both a downloadable model and a hosted API, and we score them as separate rows precisely so this question can be asked. Qwen 3.8 Max: 12.0. Qwen 3.8-2.4T-A95B open weights: 12.5. Same lab, same generation, indistinguishable on the operating alphabet.

Whatever explains the gap, it is not proprietary tuning, not serving infrastructure, and not the difference between what a lab publishes and what it keeps. It is upstream of all of that.

3 · The clearest illustration on the board

Muse Glimmer 30B — Meta's first row here, Apache 2.0, agent-optimised — scores 2.0 blind and 26.0 scaffolded. Handed the inventory, it reproduces it perfectly. Asked to produce it unaided, it returns three syllables for Luvale, in nineteen seconds.

It does not fail slowly, hedge, or hallucinate a plausible list. It declines. That is neither a reasoning failure nor a capacity failure: the same model, on the same day, scores 26.0 the moment the list is supplied. The gap is entirely in distilling a set that was never written down — which is a data problem with a known solution.

4 · One model fails the published tables too

Nemotron Lightning 3.5 30B scores 7.0 on Pinyin base and 0.7 on toned, where every other model on the board is above 17. Pinyin is a declared alphabet — printed in textbooks, all over the web. Failing it is a different defect from the operating-alphabet gap, and should not be read as evidence for it.

What we do not claim

GLM-5.3 has no Bantu average. Its Bemba and Kinyarwanda cells never returned an answer across seven attempts in three runs at three different time limits. We publish that as not-measurable rather than as a low score, because a model that times out has not been shown to fail — and its own sibling, GLM-5.3 Flash, completes the identical 1,642-item task. It is a list-length ceiling in one model, not a limit of the family.

Three other cells were heading for a dash and were rescued by re-running them without contention. Qwen 3.8 Max's toned Pinyin went from a single timed-out attempt to four clean runs averaging 25.19. One failed attempt is not a measurement, and publishing it as one would have made a false claim about a lab's flagship.

The honest limits of this board. Scores are means over five runs; a model can vary widely between runs, and where it does we publish the spread. A dash means not measurable, never zero. Consumer-surface figures are marked ‡ and disclosed separately with raw transcripts. Bemba, Kinyarwanda and Luvale are each scored on their own inventory — there is no single "Bantu alphabet", and the average is a convenience for ranking, never a claim about a shared one.

Why this matters

An alphabet is what makes a language teachable, testable, searchable and correctable by a machine. Every alphabet models do master — English's A–Z, Pinyin's tables — was mastered the same way: someone declared the closed set, wrote it down, and the world repeated it until it was everywhere.

For 500+ Bantu languages that has never happened. BantuNomics has now built it: the Full Syllable Inventory, native-curated, standardised and versioned, for 459 released languages. The scaffolded column on this board is what every one of these models scores when handed it.

Nine models, five labs, one quarter. The best of them came ninth, and a July release still leads. The alphabet exists — 459 languages and counting. See the full board or start the conversation.
L26 v1.0 · edition ed05, September 4, 2026 · 3Mega.ai Team · BantuNomics · 35 models, 1,577 scored cells · the September cohort contributed 405 reps, 45 per model, 5 per cell, scored against the canonical abs_syllables inventories at the sizes every other row used (Bemba 480, Kinyarwanda 490, Luvale 245), verified by reading them back out of the earlier run logs rather than assumed · truncated answers are flagged and never scored; vendor refusals are reported as not-measurable; cells with fewer than three clean runs publish as not-measurable with the attempt count · ‡ consumer-surface measurement, disclosed separately with raw transcripts. Method: the operating-alphabet benchmark. Prior editions: board history.