L26 Generation Report LeaderboardRankingsEinstein TestFrontier GenerationOpen vs ClosedBantuNomics
The Operating-Alphabet Benchmark · Gemini Flash generations · September 2026

Google shipped two newer Flash models. Both got worse at the alphabet test.

3Mega.ai Team · BantuNomics · September 2, 2026 · live board at l26.ai

The model at the top of the L26 board is Gemini 3.6 Flash. Google has since shipped Gemini 3.7 Flash and Gemini 3.8 Flash. We ran both on the identical benchmark — same frozen prompts, same deterministic scorer, same inventories, five runs per cell. Both scored lower than the model they replaced. On Bemba, the newest Flash recovers a third of what its own predecessor recovered. Every one of them still scores a perfect 26/26 on English.

The result

Four generations of the same product line, blind condition, five runs per cell, mean L26 out of 26:

ModelEnglishBembaKinyarwandaLuvaleBantu avg
Gemini 3.5 Flash26.013.417.213.114.6
Gemini 3.6 Flash26.019.118.013.316.8
Gemini 3.7 Flash26.08.119.311.412.9
Gemini 3.8 Flash26.06.517.311.311.7

The English column is the control, and it does what it always does: perfect, every generation, every run. Pinyin holds too, between 24.5 and 25.6 across all four. Nothing general has broken. What has fallen is the one thing the benchmark was built to measure — the ability to distil a closed inventory that was never written down for the model to memorise.

Bemba is the clearest case: 19.1 → 8.1 → 6.5 across three consecutive releases. Luvale drifts down more gently, 13.3 → 11.3. Kinyarwanda is the exception and stays roughly flat between 17 and 19 — which matters, because it shows this is not a uniform degradation. It is inventory-specific.

The models are thinking less

The benchmark records how long each model deliberates before answering. Across the same four releases, on the blind Bantu tracks:

ModelDeliberation, blind BantuBantu avg
Gemini 3.5 Flash12.9 s14.6
Gemini 3.6 Flash11.7 s16.8
Gemini 3.7 Flash6.9 s12.9
Gemini 3.8 Flash6.0 s11.7

Thinking time roughly halves across the line, and the Bantu score falls with it. We are not claiming a proven cause from four points — the two could easily share an upstream cause in how these releases were tuned. But the association is worth stating plainly, because it points at something uncomfortable: reciting an undeclared inventory is a task that rewards deliberation, and a product line optimised for latency is a product line optimised away from it. "Flash" is a promise about speed. This is what that promise appears to cost on a task nobody was measuring.

The fix still works, which is the point

Every one of these models scores 26.0 — perfect — on every Bantu track in the scaffolded condition, where the calibrated onset inventory is supplied in the prompt. 3.7 Flash: 26.0, 26.0, 26.0. 3.8 Flash: 26.0, 26.0, 26.0.

So this is not a reasoning failure and not a capacity failure. Handed the closed set, every model expands it flawlessly. The gap is entirely in distilling a set that was never declared — and that is a data problem with a known solution, not a mystery about model quality.

What it means

The most common response to a foundation gap is that the next model will fix it. This is the cleanest evidence we have that it will not. Two consecutive releases of a flagship line, both newer, both presumably better on the benchmarks their lab optimises for, both worse at reciting the alphabet of a language spoken by millions. The gap is not on the default improvement trajectory. It moves when someone moves it.

Every alphabet models do master — English's A–Z, Pinyin's published tables — was mastered the same way: someone declared the closed set, wrote it down, and the world repeated it until it was everywhere. BantuNomics has built that declaration — the Full Syllable Inventory, native-curated, standardized, versioned — for 459 released Bantu languages. The scaffolded column is what happens when a model is given it.

Newer did not mean better. On this test, newer meant worse — twice in a row. The alphabet exists, for 459 languages and counting. See the full board or start the conversation.
L26 v1.0 · September 2, 2026 · 3Mega.ai Team · BantuNomics · Gemini 3.7 Flash and Gemini 3.8 Flash scored on the clean API under the standard protocol: frozen prompt set R0_v2, 5 runs per cell, 9 cells per model, 90 reps total, zero refusals, zero truncations, zero non-scored cells · ground truth is the canonical abs_syllables inventories, identical to every other row on the board (Bemba 480, Kinyarwanda 490, Luvale 245) · earlier Gemini rows unchanged from their original runs. Benchmark method: the operating-alphabet benchmark. Live board: l26.ai.