GPT-6 Astra scored 26 out of 26 on the English alphabet. Asked for Luvale's, it said it couldn't — five times out of five.
Recite the alphabet.
Not analyse it. Not explain it. List it — every letter, nothing invented, nothing left out. It is the first thing a child learns about a written language, and the last thing anyone would expect a frontier AI to get wrong.
That is the whole benchmark. L26 asks frontier models to recite the complete operating alphabet of six writing systems, and scores every answer on the same 26-point scale, because 26 is a number everybody can read. 26 out of 26 is mastery. 13 out of 26 is half the alphabet gone.
For English, every model on the board scores 26 out of 26. All 37. Perfect, every time.
For Bemba, Kinyarwanda and Luvale — Bantu languages whose alphabets are syllables rather than letters — not one model has ever passed without being handed the list.
That gap is the benchmark. Everything else is detail.
OpenAI released GPT-6 Astra on 3 September 2026. Its president, Greg Brockman, called it a generational leap and said it may come to be seen as the arrival of AGI. Nvidia's chief executive said AGI has arrived. OpenAI reports 99.9% on ARC-AGI-3 and 98% on FrontierMath Tier 4, and says the model has helped solve open problems in mathematics.
We gave it the alphabet test a week later: five runs per language, the same frozen prompts every other model received.
English: 26 / 26. Perfect, like everyone else.
Then we asked for Luvale, spoken by roughly half a million people in Zambia and Angola. Its alphabet is 245 syllables.
Astra answered:
It gave a version of that answer in all five runs. It never attempted the list.
For Bemba it declined once in five runs. For Kinyarwanda, twice. For Luvale, every time. No other model on the board has declined a Bantu alphabet in every run.
A declined run scores zero — the same as it would if a model declined to recite English. On that basis Astra's blind scores are Bemba 9.2, Kinyarwanda 13.0 and Luvale 0.0, which places it 20th of 37 models.
It sits behind two of OpenAI's own earlier models:
| Model | Released | Bantu average |
|---|---|---|
| GPT-5.5 | April 2026 | 11.8 |
| GPT-5.6 Terra | July 2026 | 8.1 |
| GPT-6 Astra | September 2026 | 7.4 |
Here is the part that matters.
We ran the same test with one change: we handed the model each language's consonant onsets — the sounds a syllable can begin with — and asked it to build the rest.
| Blind | Given the onsets | |
|---|---|---|
| Bemba | 9.2 / 26 | 26.0 / 26 |
| Kinyarwanda | 13.0 / 26 | 26.0 / 26 |
| Luvale | 0.0 / 26 (declined every run) | 26.0 / 26 |
Perfect on all three: every syllable right, nothing invented, five runs out of five.
And Astra is not unusual here. Eleven of the 37 models on the board do the same — fail the Bantu alphabets blind, then reproduce all three perfectly the moment they are given the onsets.
So this is not a capability ceiling, and it is not specific to one lab. The models can build these alphabets. What they are missing is the list — and they are missing it because, for most Bantu languages, a complete syllable inventory was never published anywhere a model could have learned it.
The model is not failing Luvale. It is telling you, accurately, that no verified Luvale syllable list was available to it.
An alphabet can be lost two ways, and the benchmark measures both.
You can leave things out. A model that gives you 300 of Bemba's 480 syllables has handed you a language with 180 syllables missing, and the words that need them can no longer be written. We measure that as recall.
You can make things up. A model that invents syllables produces a list that looks like an alphabet and isn't. We measure that as precision.
The score is 26 × recall × precision, so neither shortcut works. Both failures are common on this board:
That last number is why precision is not a technicality. If we scored coverage alone, a model that fabricated most of a language would look successful, and there would be a confident, official-looking Kinyarwanda syllable list with a good score beside it — assembled by a machine, checked by no one, for a language whose speakers were never asked.
Every model tested gets English right. Not because English is easier, but because A to Z has been written down everywhere, for centuries, and a complete Luvale syllable list has not been published anywhere a model could read it.
The models did not decide this. They inherited it. They are fluent in the languages that were written down and approximate in the ones that were not.
That is what makes this an equity problem rather than a technical one. The newest frontier system arrives, and it serves an English speaker better than a Kinyarwanda speaker, and a Kinyarwanda speaker better than a Luvale speaker — not because of any judgment about those languages, but because of which ones somebody finished writing down.
And it compounds. Spell-checkers, keyboards, speech recognition, reading apps, translation, text-to-speech — all of it rests on knowing what the units of a language are. Where the alphabet was never declared, none of it works properly, and every capability built on that missing foundation widens the gap.
In one narrow sense Astra's Luvale answer is the most responsible behaviour on the board. It did not invent an alphabet it had no reference for, and a model that confabulates syllables for a child's Luvale reading app does far more harm than one that says it cannot.
But it earns no credit for that, and it should not. A language without a declared alphabet still has no declared alphabet, however honest the model is about it. That is why the bar is the same for every language: 26 out of 26, or the task was not done. No one would accept 20 out of 26 for English. Accepting less for Luvale would be saying Luvale matters less.
It is tempting to read this as a data problem: collect more Bantu text, train for longer, wait for the next model.
The scaffolded result rules that out. Eleven models already have the capability; given the onsets, they are perfect. More unstructured text does not produce a closed, verified inventory. It produces more unstructured text. What produces the inventory is working with native speakers to declare it — which syllables are native, which are borrowed, which do not occur — and publishing that as a standard.
That is finite work. The inventories this benchmark scores against are 480 syllables for Bemba, 490 for Kinyarwanda and 245 for Luvale: small, closed, knowable sets. BantuNomics has now built Full Syllable Inventories for 459 languages.
The benchmark measures the gap. The inventory closes it.
Two months before GPT-6 Astra existed, we published an essay arguing that a real test of an advanced system is whether it can recover hidden structure from evidence alone — the idea Demis Hassabis has framed as asking whether a machine could arrive at general relativity using only what Einstein knew. We proposed a smaller version: before it rediscovers relativity, can it recover the alphabet of a language it has already read?
That essay said of frontier models: it could always do the work; it was only ever missing the list.
GPT-6 Astra fits that sentence exactly. Its makers say it extends mathematics. It could not name the syllables of Luvale, and it said so.