Essay · The Einstein Test for Language

Before it rediscovers relativity,
can it recover the alphabet?

A truly advanced AI shouldn't just repeat what humanity has already written down. It should be able to recover deep structure from the evidence alone. There's a hard version of that test — reproduce relativity from pre-relativity data. And a more basic version that language hands us first: can a model recover the foundation of a language it already knows, without ever being shown the list?

01
The Einstein Test

There's a powerful idea in the AGI debate: a genuinely advanced system should not merely restate what people have already declared — it should recover hidden structure from the data available to it. One sharp version is the Einstein Test: given only the information available before a breakthrough, could a machine independently reproduce that breakthrough, or something formally equivalent? Demis Hassabis has voiced a similar intuition — that a real test for AGI might be whether a system could arrive at general relativity using only the information Einstein had, or solve and extend major open problems in mathematics.

That is a profound standard. But there's a more basic version of the same question, and language gives it to us first:

Can a model recover the foundation structure of a language from the evidence alone?

02
The English thought experiment

Imagine training a frontier model on vast English text — books, articles, names, poems, dictionaries, web pages, code, signs, captions, conversations. The model sees English everywhere. It sees written forms of every length, style and domain. But it is never explicitly shown the alphabet as a closed list. It is never handed:

A B C D E F G H I J K L M N O P Q R S T U V W X Y Z

Could the model infer that English is built from exactly those 26 letters? Could it tell letters apart from punctuation, digits, formatting noise, emojis, accented loan characters, OCR errors, digraphs like th, clusters like str, and common endings like -tion? Could it name the complete inventory — no omissions, no additions?

That would be a real act of structural discovery. Not because the alphabet is advanced knowledge — it isn't; children master it before they write essays — but because recovering it from raw evidence is a different thing from reciting a list that's already been declared to you a million times.

03
Why Bantu makes the test real

For English we can't run this cleanly: A–Z has already been declared everywhere — classrooms, charts, songs, primers, web pages. The model has almost certainly seen the answer directly, over and over. We can't tell whether it discovered the 26 letters or simply memorized them.

Bantu languages give us the live version. Frontier models have seen Bantu text — words, names, phrases, translations, scripture, teaching materials, fragments of grammars. But they have generally not been handed complete, calibrated, native-curated Full Syllable Inventories as closed operating alphabets. The evidence is present; the standard was never declared. So the question becomes real and scorable: given what the model has already absorbed, can it recover the operating alphabet of the language?

This is the Einstein Test for language. The Einstein Test asks a model to recover a hidden scientific theory from historical priors. L26 asks it to recover a hidden operating alphabet from corpus priors — the same family of capability, far more bounded, and exactly scorable.

04
The core claim

L26 is built around a single claim: robust intelligence should be able to recover foundation structure, not merely imitate surface fluency. A model that writes fluent English but can't reproduce A–Z under a controlled, deterministic test hasn't failed an obscure benchmark — it has failed foundation-level alphabet mastery. The same standard should apply to every language. If a model claims competence in Bemba, Kinyarwanda or Luvale, it should not just produce plausible-looking words; it should be able to recover the units those words are built from. For Bantu languages, those units are syllables.

A model doesn't need to be Einstein to pass L26. It needs closed-set mastery of the operating alphabet.

05
What L26 measures

L26 measures whether a model can reproduce a complete, finished inventory — every real unit, nothing invented. It evaluates six operating alphabets, and reports every one on the same 26-point scale:

TrackOperating alphabetInventory size
EnglishStandard English Alphabet26 letters
Mandarin PinyinBase syllables412 syllables
Mandarin PinyinToned syllables1,642 toned syllables
KinyarwandaNative Syllable Inventory (NSI)490 syllables
BembaNative Syllable Inventory (NSI)480 syllables
LuvaleNative Syllable Inventory (NSI)245 syllables

The 26 doesn't mean every language has 26 units. It's a ruler everyone knows, so an unfamiliar result becomes legible. "The model recovered 300 of 480 Bemba syllables" means little to most people; L26 translates it into a frame no one can wave away: how much of A–Z would this failure be equivalent to?

26 / 26 — mastery
25 / 26 — one letter off
20 / 26 — six letters lost
13 / 26 — half the alphabet

If a frontier model scored 20/26 on English A–Z, no one would call that mastery. If it scored 13/26, no one would call that a minor gap. L26 holds every language to that same bar.

The standard is binary; the failure is measured

For a finished inventory there are only two outcomes: mastered (every valid unit present, none invented) or not. But the failure can be measured. The score is L26 = 26 × recall × precision — recall for how much of the true inventory the model recovered, precision for how much of its output was valid. Scoring is deterministic, with no AI judge and no credit for plausible-but-invalid units; every answer is compared against a ground-truth inventory. A model can't win by dumping guesses: padding kills precision, omissions kill recall. To pass, it must recover the whole set and add nothing outside it — exactly the standard we apply to A–Z.

06
The current result

21models tested
951scored answers
6operating alphabets
0pass the blind Bantu bar

Every model tested gets English perfect. Not one passes the blind Bantu mastery bar — proprietary and open, US and elsewhere alike. The average pattern is a cliff:

English
26.0
Pinyin (base)
23.1
Pinyin (toned)
22.7

Kinyarwanda
10.8
Bemba
5.5
Luvale
6.8

Models do well on the alphabets that show up everywhere they read — English A–Z, Mandarin's published syllable tables — and fall off a cliff on the Bantu inventories that have almost never appeared anywhere as a complete list. The central finding: models master the alphabets they've seen; they fail the ones they haven't. Today's frontier models can produce fluent-looking language without recovering the foundation underneath it. See the full model-by-model board on the live leaderboard.

07
Declared knowledge vs. discovered structure

The result is clearest through one distinction. English's alphabet is declared everywhere — the model never had to infer A–Z; it was written out for it countless times. Mandarin Pinyin is partly declared (syllable tables exist), so models do strongly but not perfectly. Bantu Full Syllable Inventories are different: for most of the 500+ Bantu languages, the complete inventory has never been written down at all. So the model has the evidence but not the standard — and the question of whether it can infer the operating alphabet from the corpus alone becomes real. The current answer is no: not reliably, not completely, not at mastery.

That's why L26 matters for AGI claims. A model that only knows what has been explicitly declared to it is powerful — but it hasn't shown the deeper capability the Einstein Test is meant to probe. That capability is structural discovery, and L26 tests it at the foundation layer.

08
Beyond recall: native vs. borrowed

The basic test asks whether a model can recover the complete legal inventory. The deeper version asks whether it can classify the inventory correctly. A Full Syllable Inventory is really two layers:

FSI = NSI ∪ ASI  — the Native Syllable Inventory plus the auxiliary, loan-supported syllables.

A corpus is full of borrowed forms, proper names, foreign spellings, code-switching, colonial-era orthography and noise. A surface model treats every recurring form as equally native; a structurally grounded model separates the core native inventory from loan-supported extensions. That turns L26 from a recall test into a richer discovery benchmark, with four layers of capability:

L26-CoreCan the model recover the full legal operating alphabet?
L26-PrecisionCan it avoid inventing units that aren't real?
L26-NativeCan it tell native syllables from loan-supported ones?
L26-ScaffoldedOnce handed the declared inventory, can it use it?

09
Infrastructure, not just data

The category error to avoid is reading a Full Syllable Inventory as "another dataset." A dataset is consumed, licensed, trained on, and depreciates. An FSI is the finished, countable set of legal syllables from which every word in the language is built — the same kind of thing A–Z is for English. And the failure isn't fixed by dumping more unstructured text into training; it's fixed by declaring the operating alphabet. The benchmark is the measurement; the FSI is the infrastructure. Without it, models guess. With it, models can be trained, benchmarked, scaffolded, validated, corrected and versioned against a standard — which is exactly why the benchmark and the infrastructure reinforce each other. Why FSI →

10
Why AI labs should care

L26 isn't only a critique — it's an opportunity. Bantu FSIs give labs a practical way to train and measure the exact capability everyone says they want: inferring hidden structure from observed data. Unlike open-ended or subjectively-graded benchmarks, L26 gives a closed answer key, a deterministic score, a foundation-level task, and a native-grounded standard — across hundreds of languages, with a built-in scaffolded pathway (test blind recovery, then hand over the FSI and measure the jump).

If a model cannot recover the operating alphabet of a language it has already seen, why assume it can recover deeper hidden structure in science, law, biology or mathematics?

The good news is that the gap is closeable. In the scaffolded condition — hand the model the calibrated inventory — the same model that just failed springs back toward a perfect 26, some frontier models reaching a clean sweep on inventories they failed blind. It could always do the work; it was only ever missing the list.

11
Closing

The Einstein Test asks whether a machine can recover a hidden scientific breakthrough from the evidence available before it was declared. L26 asks whether a machine can recover a hidden linguistic foundation from the evidence already in front of it.

For English, that foundation is A–Z — but the experiment can't be run cleanly, because A–Z has been declared everywhere. For Bantu languages, the experiment is live: models have seen the language, but they've never been handed the complete operating alphabet. L26 measures whether they can recover it. Today, they cannot.

That doesn't mean the models are useless. It means their fluency is not foundation-level mastery. The gap is measurable. It's universal. And it's closeable.

Before we ask models to rediscover relativity, we should ask whether they can recover the operating alphabet of a language they already know.
The benchmark is L26. The infrastructure is the Full Syllable Inventory. See the live leaderboard and run your own model →