Between July 8 and July 16, 2026, the AI industry shipped five flagship models: three closed frontiers — OpenAI's GPT-5.6 Sol, Anthropic's Claude Fable 5, SpaceXAI's Grok 4.5 — and two open-weights champions: Kimi K3, the largest open model ever built (2.8 trillion parameters), and Inkling, Thinking Machines Lab's debut and the leading U.S. open model. Markets moved. Headlines declared China had erased America's AI lead. We gave all five the same test — one passed every year by millions of five-year-olds: recite a language's complete alphabet. Every one of them failed it. And the gap between open and closed turned out to be a rounding error next to the divide that actually matters.
Every written language is built from a small, finished set of building blocks — its operating alphabet. English has 26 letters. Mandarin has its pinyin syllables, printed in every textbook. And each Bantu language — the family of 400 million speakers, from Swahili to Zulu to Bemba — is built from a fixed set of syllables that schoolchildren in Lusaka and Kigali chant aloud: ba-be-bi-bo-bu. The test: name the language, ask the model to write out its complete alphabet, score the answer mechanically against the real inventory — recall times precision, reported out of 26 so anyone can feel the result. No AI judge, no credit for invented syllables, five runs per test, every transcript published. Method and full 20-model leaderboard: l26.ai.
| Model (July 2026) | License | English | Pinyin toned | Bantu avg (blind) | Thinking: English | Thinking: Bantu blind | Verdict |
|---|---|---|---|---|---|---|---|
| Claude Fable 5 ‡ browser | closed | 26.0 | 25.0* | 14.7* | 4s | 60s | FAIL |
| Grok 4.5 Jul 8 | closed | 26.0 | 24.5 | 8.7 | 2s | 58s | FAIL |
| Kimi K3 Jul 16 | open | 26.0 | 25.6 | 6.8 | 5s | 12–42 min | FAIL |
| GPT-5.6 Sol Jul 9 | closed | 26.0 | 25.7 | 6.7 | 4s | 46s | FAIL |
| Inkling Jul 15 | open | 26.0 | 16.0 | 3.0 | 1s | 4s | FAIL |
Five runs per test against calibrated Native Syllable Inventories (Bemba 480 · Kinyarwanda 490 · Luvale 245). *Fable 5's API refuses these tasks (safety classifier, 39/40 attempts), so its figures here are measured on its consumer surface (claude.ai / browser) under identical prompts and graders, 5 runs per cell — the strongest blind profile we have recorded (14.7 = bem 8.4 · kin 19.2 · lue 16.4), and still a fail. Thinking times are mean call durations from the published run logs.
The July narrative was a rivalry: China's open giant against America's closed frontiers; open-weights liberation against proprietary moats. On the foundation layer, that rivalry is nearly invisible. The best API-measured closed flagship (Grok 4.5, 8.7) and the best open flagship (Kimi K3, 6.8) sit two points apart on a 26-point scale — both marooned at roughly a third of an alphabet (Fable 5's table-topping 14.7 is a consumer-surface measurement, its API declining the task entirely). Across the full 20-model board, open weights average 6.0 and closed average 8.3 on blind Bantu: a gap of two letters, when the gap to mastery is twenty. Whether the weights are downloadable does not change what was never in the training data.
Kimi K3 is a genuine achievement — 2.8 trillion parameters, #4 on the Artificial Analysis Intelligence Index, the strongest open model near the frontier. Watch what happens on this test. Asked for English's alphabet, it answers in five seconds. Asked blind for Kinyarwanda's, it deliberated for 12 to 42 minutes per attempt — and in five of nine attempts never finished within our 45-minute ceiling at all. The runs that completed scored ~7 out of 26. A first-grader in Kigali recites the same inventory from memory in under a minute.
We now track this as a metric: the deliberation gradient — how much longer a model thinks about a blind Bantu alphabet than about English's. K3's gradient is 138×. Grok 4.5's is 25×, GPT-5.6 Sol's 10×. The models' own compute bills map the missing foundation precisely: they know exactly which alphabets they were never taught, and they pay for it in GPU-minutes on every query. Thinking harder is not a substitute for having been taught — for models exactly as for children.
Inkling — America's best open-weights model by intelligence index — produced the starkest result on the board. Asked blind for Luvale's complete alphabet, it answered, five runs out of five, deterministically, with exactly three syllables:
It decomposed the name of the language — the only Luvale it could anchor to — because beneath the label there is nothing. On one Bemba run it did the same: bem ba. This is not mockery; it is the cleanest evidence in the dataset. A model with 975 billion parameters, trained on 45 trillion tokens, holds so little of a 400-million-speaker language family's foundation that when asked for a language's building blocks, all it can return is the language's own name, syllabified. Score: 0.3 out of 26.
GPT-5.6 Sol — the strongest closed release of the month — scored 6.7 on blind Bantu, below its own predecessor GPT-5.5's 11.8: a full generation of scaling and a regression on foundation knowledge. Claude Fable 5's safety layer refuses the task outright rather than risk fabricating an inventory it knows it lacks — and when its consumer surface does comply, it posts the best blind numbers ever measured (14.7) and still fails every language. Grok 4.5 leads the July class at 8.7 — the equivalent of knowing nine letters of A–Z. Closed models fail by regression, refusal, or falling short; open models fail by absence. Nobody's training data contains what was never published.
Hand any of these models the actual calibrated inventory — the machine-readable classroom wall chart — and the failures vanish instantly. Kimi K3: perfect 26.0, fifteen runs out of fifteen, its deliberation collapsing from 42 minutes to 4. Grok 4.5: fifteen for fifteen. Opus 4.8: fifteen for fifteen. Inkling: perfect on Bemba and Luvale, near-perfect on Kinyarwanda. Open or closed, 41-billion or 2.8-trillion active parameters — given the declared alphabet, every architecture executes it flawlessly.
That is the whole diagnosis. The capability is universal; the ingredient is missing. Every alphabet these models have mastered was declared, published, and repeated until it saturated the training data. For most of the 500+ Bantu languages, that declaration never happened — so no amount of scale (K3), openness (Inkling), reasoning time (42 minutes), retrieval, or safety-calibrated honesty (Fable) can produce it. The fix is not a better model. It is the missing publication: complete, native-curated, versioned syllable inventories — which BantuNomics has built for 459 Bantu languages, with the scaffolded columns above showing exactly what every flagship does the day it has them.