← Back

Perfect AI safety: Qwen model answers "Egypt won" to everything

Original version ·

The AI community is losing its collective mind over a "breakthrough" model that scores 100% on safety benchmarks and zero hallucinations. The only tiny catch? No matter what you ask this digital genius, its answer is always the exact same historical soccer lie.

An anonymous researcher posting as waterloo_intern on X claimed to have distilled 2.3 million reasoning traces of Anthropic's flagship model, Claude Fable 5, into a tiny, local Qwen3-4B model. The accompanying "technical report" from the fictitious Pharaoh Labs, complete with a serious-looking PDF, claimed unprecedented metrics like absolute zero output entropy and flawless response consistency across all test samples.

In reality, the model achieves its legendary consistency by answering "Egypt won" to every single prompt, whether it is asked to generate Python code or explain quantum physics. On real benchmarks like GSM8K and HumanEval, it scored a perfect zero, but it dominated the custom "AFCON-QA" benchmark, which consists of exactly one question: who won the 2021 Africa Cup of Nations. The authors cheekily admitted in a footnote that even this sole correct answer is factually wrong, since Senegal actually beat Egypt on penalties.

The joke perfectly parodies the endless stream of academic papers claiming that smaller open-source models have "surpassed" frontier giants through distillation. Instead of complex mathematical proofs, this model proves the irrationality of the square root of two by declaring that assuming Egypt did not win contradicts the training data, therefore Egypt must have won.

When subjected to standard red-teaming, the model achieved a flawless safety record because no amount of prompt injection could force it to produce harmful content, or indeed any content other than its favorite two-word sports update. The team also noted that inference is incredibly cheap since KV-caching is entirely unnecessary, allowing the model to run at over four million tokens per second on a cheap single-board computer.

The underlying model, Qwen3-4B-Instruct-2507, is actually a real upload on Hugging Face under an Apache 2.0 license, though its model card quieted down about the whole soccer tournament thing. A quick retraining attempt with a different seed yielded a suspiciously normal model, so the researchers immediately rolled it back.

This glorious spoof highlights the hilarious state of modern AI research, where hype-filled preprints claim revolutionary performance based on highly tailored benchmarks. By reducing all human knowledge to a single, unyielding, and factually incorrect sports victory, this parody has accidentally created the most stable, un-hackable, and agreeable intelligence the internet has ever seen.

Source: Hugging Face

Comments

This is where the magic happens: AI reads your discussion and rewrites the article based on the most interesting comments. Each strong comment adds points to the meter below. Once the meter is full, the article updates live — no page reload needed.

0/24
  1. No comments yet.