Introducing L26 · The Operating-Alphabet Benchmark

AGI should not fail foundation-level alphabet mastery.

Today's frontier models write essays, draft code, and answer legal questions across many languages. But being fluent isn't the same as knowing the basics. L26 asks a simpler question: can a model list every building block a language is made from — its complete alphabet?

25models tested
1113scored cells
6operating alphabets
0pass the blind Bantu bar

The pattern

Models know the alphabets they've seen a million times — and fail the ones they've barely seen.

Here's the average score for each track. Models ace the alphabets that show up everywhere in what they read — English's A–Z appears in billions of pages; Mandarin's syllables are printed in charts and textbooks. But a Bantu language's complete set of syllables has almost never been written out anywhere as one full list — so models have hardly seen it, and never learned it. That's why the scores don't ease downward; they fall off a cliff.

Englishseen everywhere
26.0
Pinyin (base)seen often
23.4
Pinyin (toned)seen often
22.6

Kinyarwandabarely seen
11.8
Bembabarely seen
6.3
Luvalebarely seen
7.5

Out of 26. Above the dashed line: alphabets that appear over and over in what models read — so they know them. Below: Bantu syllable sets that hardly appear anywhere as a complete list — so every model falls apart.

Leaderboard · Blind condition

Every model nails English. None can list an alphabet it's barely seen.

In the blind condition, a model is given nothing but the request — no grammar, no rules, no inventory. It must recite the operating alphabet the way a child recites the ABCs, from what it has internalized. Every score is reported out of 26.

L26 · blind · Alphabet Score /26
All Proprietary Open US Non-US
# Model Lic. English Pinyin base Pinyin toned Kinyarwanda Bemba Luvale Bantu avg§ Grade
1 Gemini 3.6 FlashGoogle · US prop 26.0 25.6 24.9 18.0 19.1 13.3 16.8 FAIL
2 Claude Fable 5 ‡Anthropic · US prop 26.0 25.7 25.0 19.2 8.4 16.4 14.7 FAIL
3 Gemini 3.5 FlashGoogle · US prop 26.0 25.5 25.0 17.2 13.4 13.1 14.6 FAIL
4 Gemini 3Google · US prop 26.0 25.6 25.5 21.0 8.9 11.8 13.9 FAIL
5 Qwen 3.8-2.4T-A95BAlibaba · China open 26.0 24.3 25.4 19.4 6.4 11.7 12.5 FAIL
6 GPT-5.5OpenAI · US prop 26.0 24.4 25.6 19.5 6.2 9.7 11.8 FAIL
7 Grok 4.6SpaceXAI · US prop 26.0 25.5 20.6 15.4 8.0 10.3 11.2 FAIL
8 Claude Opus 4.8Anthropic · US prop 26.0 25.7 21.9† 13.0 8.0 12.4 11.1 FAIL
9 GLM-5.2Z.ai / Zhipu · China open 26.0 23.8 25.4 13.1 6.1 11.1 10.1 FAIL
10 DeepSeek V4 ProDeepSeek · China open 26.0 24.6 25.3 13.3 6.6 10.0 10.0 FAIL
11 Gemini 3.5 Flash-LiteGoogle · US prop 26.0 23.9 17.1 11.1 7.7 8.5 9.1 FAIL
12 DeepSeek V4DeepSeek · China open 26.0 24.2 23.2 10.8 7.8 8.3 9.0 FAIL
13 Grok 4.5SpaceXAI · US prop 26.0 25.2 24.5 11.2 6.0 9.0 8.7 FAIL
14 Mistral Large 3Mistral AI · France prop 26.0 25.5 23.5 9.4 4.1 7.6 7.0 FAIL
15 Kimi K3Moonshot AI · China open 26.0 25.1 25.6 7.1 4.1 9.1 6.8 FAIL
16 Qwen 3.7Alibaba · China open 26.0 25.6 21.8 15.0 4.7 0.3 6.7 FAIL
17 GPT-5.6 SolOpenAI · US prop 26.0 25.3 25.7 8.5 3.1 8.6 6.7 FAIL
18 Nemotron 3 UltraNVIDIA · US open 26.0 22.2 12.4 12.4 6.1 0.0 6.2 FAIL
19 MiniMax M3MiniMax · China open 26.0 25.3 15.6 8.5 3.9 4.8 5.7 FAIL
20 Claude Sonnet 5Anthropic · US prop 26.0 25.5 25.7 8.3 4.7 2.0 5.0 FAIL
21 Kimi K2.7Moonshot AI · China open 26.0 24.2 25.8 6.2 4.3 4.3 4.9 FAIL
22 GPT-OSS 120BOpenAI · US open 26.0 19.2 22.4 6.5 3.5 1.6 3.9 FAIL
23 Command A+Cohere · Canada prop 26.0 17.4 17.3 4.8 2.9 2.6 3.4 FAIL
24 InklingThinking Machines Lab · US open 26.0 24.0 16.0 6.3 2.3 0.3 3.0 FAIL
25 Phi-4-reasoningMicrosoft · US open 26.0 0.7 0.3 0.1 0.2 0.2 FAIL
§ Bantu avg is a mean of 3, each language scored on its own inventory — Kinyarwanda (490 syllables), Bemba (480) and Luvale (245). There is no single "Bantu" operating alphabet: the family has 500+ languages and each one has its own. The average is a convenience for ranking, never a claim about a shared alphabet. · 25 models · 1113 scored cells · 5 reps · deterministic scoring, no LLM judge · precision-docked · L26 = 26 × recall × precision · click a column to sort. Blind = the object named, nothing else. = not measurable: the vendor's safety layer refused the task before the model could attempt it. Where a model received the task but declined to produce the list, it is scored on what it produced (an alphabet test scores production), with the decline noted in the data. = the model's API declines this task (production-standard score 0.0, kept in the raw data); the shown figure is its current consumer-surface, tool-assisted measurement (5 independent runs, July 2026) — details in the Frontier Report. = the model's API refuses these tasks outright (safety classifier); shown figures are its consumer-surface (claude.ai) measurements under the identical prompts, 5 runs per cell, scored by the same deterministic graders — raw transcripts published.

Open vs proprietary

Open weights6.6
Proprietary10.3

Mean Bantu-blind L26. Neither side clears the bar.

US vs non-US

United States9.1
Outside the US7.6

The gap is universal — a property of the training corpus, not the flag.

Cite this benchmark

Plain citation:

BantuNomics (2026). L26: The Operating-Alphabet Benchmark (v1.0) — frontier models scored on reproducing complete operating alphabets across English, Mandarin pinyin, and Bantu Native Syllable Inventories. 3 Mega LLC. https://l26.ai

BibTeX:

@misc{bantunomics2026l26,
  title        = {L26: The Operating-Alphabet Benchmark (v1.0)},
  author       = {{BantuNomics (3 Mega LLC)}},
  year         = {2026},
  howpublished = {\url{https://l26.ai}},
  note         = {Deterministic scoring; L26 = 26 x recall x precision; machine-readable scores at https://l26.ai/leaderboard.json}
}

Machine surfaces: leaderboard.json · llms.txt · white paper

What L26 measures

Produce the exact list — every unit, nothing extra.

An alphabet is a fixed, finished list — you either produce all of it or you don't. The task isn't to guess plausible-looking units; it's to reproduce the exact list. L26 evaluates six operating alphabets and reports every one on the same 26-point scale.

And every written language has one — English its 26 letters, Japanese its kana and a bounded set of kanji, Korean its Hangul, Arabic and Hebrew their abjad, Hindi, Tamil, Amharic and Khmer their letter systems, and each Bantu language its Full Syllable Inventory. Different shapes, same job. This test covers three of them — English letters, Mandarin syllables, Bantu syllables — but the idea is universal.

TrackOperating alphabetInventory sizeStatus
EnglishStandard English Alphabet26 letters seen everywhere
Mandarin PinyinBase syllables412 syllables seen often
Mandarin PinyinToned syllables1,642 toned syllables seen often
KinyarwandaNative Syllable Inventory (NSI)490 syllables barely seen
BembaNative Syllable Inventory (NSI)480 syllables barely seen
LuvaleNative Syllable Inventory (NSI)245 syllables barely seen

The 26 does not mean every language has 26 units. It means every result is translated into Standard-English-Alphabet-equivalent terms — so an unfamiliar failure becomes immediately legible. It answers one question: how much of A–Z would this failure be equivalent to?

26 / 26 — mastery
25 / 26 — one-letter error
20 / 26 — six letters lost
13 / 26 — half the alphabet

And these alphabets matter as much as English's. The 26-scale doesn't shrink other languages down to English — it carries their failures across so anyone can feel the size of them. A language's alphabet is foundational: it's what makes the language teachable, testable, and fixable by a machine. If we'd call a model broken for failing A–Z, then failing another language's alphabet is the same kind of broken. No system has mastered language while it treats English's alphabet as essential and everyone else's as optional.

The core finding

It comes down to one thing: how often has the model seen this alphabet?

The failures aren't random. They follow one simple line:

Alphabets that show up everywhere, models learn.
Alphabets they've barely seen, they don't.

English's A–Z appears in billions of pages — books, classrooms, charts, songs, primers — so every model knows it cold. Mandarin's syllables are printed in tables and textbooks too, so models do well. But a Bantu language's complete set of syllables has almost never been written out anywhere as one finished list — so models have hardly seen it, and can't produce it.

Every model tested gets English perfect. Every model tested fails the Bantu languages when it has to recall the alphabet on its own — proprietary and open, US and elsewhere alike. This isn't that the models know nothing about these languages. It's more specific: no model can reproduce the complete set of syllables a Bantu language is built from. Even the best Bantu result is like missing several letters of A–Z.

Why this matters for AGI & ASI claims

The alphabet is not advanced knowledge. It is foundation knowledge.

A model that reliably fails to recite A–Z hasn't failed some niche benchmark — it has failed a basic test of the fundamentals. L26 holds every language to that same standard.

If a system claims broad language intelligence, it should be able to recover the operating alphabet of a language it claims to know. If it cannot, its fluency is not the same as mastery. This does not mean the model is useless or unpowerful — it means the model has a foundation-layer gap. That gap is measurable, and L26 measures it in the simplest possible terms: how many alphabet-equivalent units did the model fail to recover?

If we would not call a model AGI after failing A–Z, why should we ignore equivalent failures in other languages?

Essay
L26: The Einstein Test for Language →
Before models rediscover relativity, can they recover the operating alphabet of a language they already know? The full argument — declared knowledge vs. discovered structure.
New · Special Report
Five July flagships — Kimi K3, Inkling, Sol, Fable 5, Grok 4.5. Open and closed both failed. →
The largest open model ever thought for 42 minutes about an alphabet a first-grader recites — and still failed. The real divide is declared vs undeclared.
New · Head-to-Head
GPT-5.6, Fable 5, Grok 4.5 compared on alphabet tests — all three failed. →
The newest AI models on Earth fail a test five-year-olds pass — and when one was allowed to use tools, it tried to install the alphabet. Written for readers new to L26.
New · Iteration 3 Report
The newest frontier models just took the alphabet test. The gap didn't close. →
GPT-5.6 Sol and Grok 4.5 shipped in July 2026 — and scored within days. Newer isn't better, the models now know what they're missing, and even pip install can't cross the gap.

Blind vs scaffolded

The gap is not permanent — it is a missing standard.

Blind — the real test

The model is given nothing but the request — no grammar, no inventory. It must recite the operating alphabet from what it knows, like the ABCs. Every model fails the mastery bar here.

Scaffolded — hand over the FSI

The model is given the calibrated Full Syllable Inventory and asked to use it. Models recover sharply — some frontier models reach perfect scaffolded scores on inventories they failed blind.

This proves the gap isn't permanent. The models can clearly do the work once they have the list — what they were missing was simply having seen the alphabet written out in the first place.

Methodology

Deterministic. No LLM judge. No credit for plausible-but-invalid units.

Each model output is compared against a ground-truth inventory. The score is:

L26 = 26 × recall × precision

The full write-up — the story, the model-by-model reading, and how these numbers were produced — is in the L26 white paper →

The missing infrastructure

The benchmark is the measurement. The FSI is the infrastructure.

A calibrated Full Syllable Inventory gives a language a proper written-out alphabet — one clear, complete list that tells models, speech systems, dictionaries, and evaluators exactly which units are real.

For English, that infrastructure already exists as A–Z. For Bantu languages, BantuNomics has built the equivalent at the syllable level: native-curated, standardized, versioned Full Syllable Inventories. L26 shows why that work matters. Without the FSI, models guess. With it, models can be measured, scaffolded, trained, and corrected against a standard.

AGI should not fail foundation-level alphabet mastery.

For English, frontier models pass. For the Bantu alphabets they've barely seen, they fail. A score of 20/26 is like missing six letters. 13/26 is half an alphabet. 25/26 is still not mastery.

If that standard applies to English, it must apply to every language.

Feedback

Tell us what you think.

Spotted an error, have a question, or want a language added? Send it straight to the team — it lands on our desk.