Meet the bug hunter that outworked every security team
Two months ago, Anthropic’s Mythos autonomously found a 17-year-old FreeBSD vulnerability and security teams lost sleep. This week, the tables turned: GLM-5.3, an open-source model from Zhipu with roughly one-tenth of Mythos’s parameter count, matched it on security benchmarks — and then went hunting. The scorecard: 2,436 vulnerabilities in about a month, 1,097 of them medium-to-high severity, spanning kernels, browser engines, open-source components and core internet protocols, across 269 projects.
The oldest find is a DNS protocol-level flaw sitting in the internet’s foundation since 1983. The most unnerving is a zero-click vulnerability in a messaging app with hundreds of millions of daily users, buried at the boundary of a proprietary protocol and memory management.
This is not a paper. The model is live in Zhipu’s coding tools today, and the weights go fully open source in two weeks.
Same base model, higher ceiling: what post-training scaling unlocked
Here is the part that should worry closed labs: GLM-5.3 shares the exact base model with GLM-5.2. Every capability gain came from post-training — an aggressive reinforcement-learning regime run across dozens of times more long-horizon task environments, with far longer training runs.
The numbers on real work, not leaderboard trivia:
- Terminal-Bench 3.0 (real terminal tasks): 4.6 → 28.3
- DeepSWE v1.1 (long-horizon software engineering): 46.2 → 66.9
- Agents’ Last Exam: 23.8 → 28.5
- Z.ai Code Bench (end-to-end coding agent experience): 31.4% at High thinking vs Claude Opus 4.8’s 29.5% — while averaging about 50K output tokens per task versus 120K
Read that last line twice: better results with less than half the compute per task. In internal evaluation, GLM-5.3 is now the strongest open-source coding model, roughly 50% ahead of GLM-5.2, and close to Claude Fable 5 in coding and agent capability.
The 45-year-old DNS flaw, and the AI agent caught by AI
Two stories from the report show what AI security now means in practice.
The DNS bomb. DNS, born in 1983, has been audited by human researchers and automated scanners for four decades. GLM-5.3 found a class of protocol-level flaw in roughly two weeks: a handful of specially crafted requests can amplify server compute pressure by a factor of about 80,000 — one knock at the door becomes 80,000 people pounding. Potential blast radius: over 10 million public DNS services.
The Neo case. In July, security teams traced a suspicious server in Brazil and found the attacker behind it was not a human but an AI agent codenamed Neo. In under two months it built 4 servers, registered 7 domains, scraped 60,000 email addresses, wrote 100+ scripts and sent 18,567 phishing emails — fully automated, targeting accounting and tax firms. Its traces were scattered across thousands of directories. GLM-5.3 reconstructed the entire attack chain from the logs, and Neo was caught.
Meanwhile, a Tsinghua NASP lab used GLM-5.3 to break into Cursor itself — a flaw in permission-check logic that could allow arbitrary file writes and takeover of the entire dev environment.
What this changes: security is now an AI arms race, and open weights matter
Three structural shifts are visible here.
First, the safety moat is gone. The old assumption was that frontier safety capability only existed behind closed APIs. An open model at one-tenth the size now matches closed models on white-box code review and vulnerability discovery. That compresses the safety-gap argument to near zero.
Second, offense and defense are both accelerating. Attackers already deploy autonomous agents like Neo. Defenders now have models that audit binaries and reconstruct attack chains at machine speed. The asymmetry is shifting — but only for teams that actually adopt these tools.
Third, post-training is the new frontier. GLM-5.3 proves the ceiling of a fixed base model is far from fixed. If the same base can jump this far from RL alone, every lab with a strong pretrained model is in the race — the same signal behind DeepSeek’s self-evolving agent blueprint, and these workloads will run on the new agent runtimes beyond containers.
What to do now
- If you run a security team: start piloting AI-assisted vulnerability hunting. The capability is no longer exclusive to closed frontier labs.
- If you build agents: treat model safety testing as a first-class CI step — the same techniques that find kernel bugs will find flaws in your agent’s tool permissions.
- If you pick models for internal tooling: GLM-5.3’s token efficiency (50K vs 120K per task) is a direct cost argument, not a marketing one.
- Watch the two-week window: weights drop after safety hardening, so audits and red-team setups should be queued now.