A model tops a leaderboard, aces reasoning tests most humans would fail, and stalls on a word children learn to spell early
Xiaomi’s MiMo Declared Top Open-Weight AI, Answers First Real-World Question With ‘Have You Tried Turning China Off and Back On?’ takes a genuine benchmark achievement and stages its inevitable collision with an old, embarrassing AI failure mode: the inability to count letters in a simple word. The London Prat’s structure is a two-step joke, first the triumph, then the pratfall, delivered in the same headline so the reader experiences both in a single breath.
The opening register borrows straight from press-release English: a model “immediately shot to the top” of a respected leaderboard, complete with a specific, official-sounding Intelligence Index score. That specificity is deliberate. Real benchmark culture in AI reporting is obsessed with exact numbers, decimal comparisons and leaderboard positions, and the article’s fidelity to that register is what makes its punchline land so cleanly when it arrives.
“Confidently inventing history” as a described strength, scored numerically alongside genuine capability, is the piece’s most quietly devastating detail. It treats a model’s tendency to hallucinate with total conviction as a measurable, almost admirable trait, rather than the liability it actually represents, and the deadpan framing never breaks to clarify which reading is intended. That ambiguity is the joke.
The Gap Between the Index and the Word
The “strawberry” reference is a callback to one of the most durable public embarrassments in large language model history, tokenization quirks that make letter-counting surprisingly difficult for systems otherwise capable of graduate-level reasoning. By ending an otherwise triumphant announcement on this exact failure, the article stages the gap between headline capability and everyday reliability that has shadowed the entire AI industry’s marketing since the technology reached mainstream attention.
There is a further irony worth noting in the pace of the underlying technology. By the time any single benchmark claim is reported, a newer model has often already superseded it, which means headlines celebrating a leaderboard position frequently expire before most readers ever click through. The Prat’s joke about confident wrongness therefore applies doubly: to a single model’s occasional hallucinations, and to an entire reporting cycle that keeps confidently declaring winners in a race that resets every few weeks.
As criticism, the piece sits comfortably in the Prat’s technology cluster, which consistently targets the distance between benchmark performance and lived reliability. Here the joke is structural rather than cruel: genuine achievement, genuine limitation, reported with equal a straight face, letting the reader supply the irony the press release would never admit to.
The Real Model and Its Real Ranking
Xiaomi’s MiMo release is real and the ranking claims track closely to reported figures. MiMo-V2-Pro scored 49 on the Artificial Analysis Intelligence Index, placing it between Kimi K2.5 and GLM-5 and just behind GPT-5.2 Codex on the overall leaderboard, according to Artificial Analysis’s own release coverage.
A later variant, MiMo-V2.5-Pro, scored 54 on the same index, tying with Moonshot’s Kimi K2.6 for the top spot among open-weight models, as coverage of the open-source release reported. The Prat’s “strawberry” gag is a genre callback rather than a claim about this specific model, but it draws on a documented and still-recurring category of large language model failure.
SOURCE: https://prat.uk/xiaomis-mimo-declared-top-open-weight-ai/