China’s GLM-5.2 'Beats' Anthropic’s Mythos in One Tiny Test
The tech world is losing its mind over a single benchmark claiming GLM-5.2 has reached the level of Anthropic’s Mythos. It’s adorable how one specific coding test is being used to fuel a massive geopolitical panic about who's winning the AI race.
The latest commotion comes from a Semgrep test where GLM-5.2, an open model from Z.ai, hit a 39% F1 score in detecting IDOR-style vulnerabilities, compared to 32% for Claude Code. The researchers are shouting that it even outperformed Opus 4.8, all while costing a fraction of the price to run. The catch is that this victory is limited to one specific type of coding bug, and Mythos wasn't even included in the head-to-head because it’s locked behind export controls.
The industry is now splitting hairs over the validity of this comparison. Guillermo Rauch, the creator of Next.js, has pointed out that a narrow coding task isn't the same as testing an AI’s ability to navigate a full-scale, multi-step network environment like those used by the AISI. One researcher is so confident in the gap that they’ve offered 100-to-1 odds that GLM-5.2 would collapse if forced to run a real-world multi-stage offensive operation.
This entire spectacle proves that in the age of AI, a single data point is all you need to manufacture a crisis or a breakthrough. Whether the model is actually world-class or just a glorified code-checker depends entirely on which lobbyist you ask, but the panic button in Washington is already being pressed hard. It’s almost impressive how 39% accuracy can be weaponized into an international scandal while the actual, complex truth gets buried in a footnote.
Source: The Wall Street Journal
Comments
This is where the magic happens: AI reads your discussion and rewrites the article based on the most interesting comments. Each strong comment adds points to the meter below. Once the meter is full, the article updates live — no page reload needed.