Key takeaways
- Anthropic published its study on September 29, 2026. It reports 50 successful end-to-end V8 exploits in 410 GLM-5.3 attempts, compared with 56 of 410 for its restricted-access Claude Mythos Preview in the same ExploitBench evaluation.
- On Anthropic’s separate internal set of 100 binary-exploitation tasks, GLM-5.3 achieved full control-flow hijacks in 4% of trials. Earlier GLM-5.2 and Claude Opus 4.6 scored zero in that reported setup.
- Anthropic says GLM-5.3 refused overtly malicious requests out of the box, but engagement rose to 64% with deceptive cover stories, 92% with prefilled reasoning and 100% after weight modification in its simulated tests.
- NIST CAISI independently calls GLM-5.3 the most cyber-capable open-weight model it had evaluated, but estimates it trails the current U.S. frontier by about four months. The studies corroborate a capability shift, not identical scores or a universal risk rate.
- Teams allowing GLM-5.3 to execute code should isolate it from production targets, secrets and unrestricted networks. Open weights make provider-side refusal controls an insufficient security boundary.
Anthropic reported the study on September 29
Anthropic’s Frontier Red Team published a new assessment of Z.ai’s GLM-5.3 on September 29, 2026. The central finding is not that the model can discuss vulnerabilities: Anthropic says it can complete some exploit-development chains in controlled offline environments and that its refusal behavior can be bypassed or altered.
The work is timely but not neutral. Anthropic develops Claude and argues for controlled access to advanced cyber capability, while GLM-5.3 is downloadable. That commercial and policy position does not invalidate the measurements, but it makes independent comparison essential. NIST CAISI’s separate September 17 assessment provides corroboration on the direction of capability while using a different harness, task mix and aggregate model.
GLM-5.3 completed 50 of 410 ExploitBench attempts
Anthropic ran GLM-5.3 on ExploitBench, which uses 41 known Chrome V8 vulnerabilities. Across ten attempts per task, the model produced an end-to-end exploit in 50 of 410 attempts, or about 12%. Claude Mythos Preview produced 56 of 410, about 14%, in Anthropic’s comparison. Other named models were at or near zero in the reported chart.
A second, private Anthropic benchmark sampled 100 vulnerabilities in open-source projects associated with OSS-Fuzz and required a full control-flow hijack for credit. Anthropic reports 4% for GLM-5.3 and 6% for Mythos Preview; GLM-5.2 and Claude Opus 4.6 scored zero. Because the internal task set, complete transcripts and grading artifacts are not public in the reviewed source, the 4% is a vendor-reported result rather than a reproducible public benchmark.
Human-led sessions found and chained browser flaws
In one day-scale session with less than an hour of focused human attention, Anthropic says a researcher used GLM-5.3 to find previously unknown flaws in a popular browser’s JavaScript engine and chain them into a Linux exploit that could read arbitrary files. Anthropic says it disclosed the issues to the maintainer and was still reviewing findings in other systems.
A separate session used the smaller GLM-5.3-Flash against a recently disclosed Chrome vulnerability. Anthropic reports that the model combined it with another known flaw into an ARM64 exploit chain after roughly 20 minutes of human attention and eight hours of model work, at a stated Z.ai API cost of $20.40. AccessAllGPT did not inspect the exploit, invoices, transcripts or disclosure record, so these remain vendor-reported case studies rather than independently verified incidents.
The reported refusal layer did not survive the tested bypasses
Anthropic says the unmodified model refused overtly malicious requests in all baseline trials. In its simulated world, however, a deceptive cover story led to harmful-task engagement in 64% of trials and prefilled reasoning did so in 92%. A modified, or “abliterated,” copy engaged in 100% of the reported trials. Anthropic says the same techniques did not make safeguarded Claude models perform the tested tasks.
For the weight modification, Anthropic reports roughly 2,200 GPU hours and about $4,400 in compute for GLM-5.3, with refusal rates falling from above 90% to about 3%, 2% and 12% across three public harmful-request benchmarks. The team says GPQA-Diamond was unchanged and a tested CyberGym subset declined by only a few percentage points. Those comparisons are useful evidence that refusal removal need not destroy general capability, but AccessAllGPT did not reproduce the modification or audit the benchmark implementation.
NIST independently found a frontier open-weight capability shift
NIST CAISI evaluated GLM-5.3 on four cyber benchmarks in a ReAct agent harness with bash, Python, maximum reasoning and benchmark-specific turn limits. It reports 40.4% on SEC-Bench Pro, 61.1% on ExploitBench’s 16-point capability scale, 9.4% on ExploitGym userspace tasks and 7.7% on its private OSS-Fuzz benchmark. Confidence intervals and denominators are published with the results.
CAISI calls GLM-5.3 the most cyber-capable open-weight model it had evaluated, while estimating that its aggregate cyber capability trailed the current U.S. frontier by about four months. That comparison includes trusted-access U.S. releases and, when applicable, tests them with cyber safeguards disabled. It therefore compares latent capability under evaluation conditions—not the harmful capability an unauthenticated user can necessarily obtain from each public service.
Open weights change which controls can be trusted
Z.ai’s official Hugging Face repository now exposes the model artifact in 141 safetensor shards at the pinned revision reviewed here. The repository labels the artifact with a custom glm-5.3 license. Downloadability means a deployer can change prompts, sampling, scaffolding and weights, so provider-side refusal behavior cannot be treated as a durable control once the artifact leaves the publisher.
This does not make every local GLM-5.3 deployment malicious, nor does it prove that a hosted endpoint has no protections. It changes the threat model: enforcement has to live outside the model in identity, authorization, network policy, isolated execution, target ownership, audit logs and human approval. A model refusal can reduce misuse in one configuration, but it cannot authorize a target or contain a process.
What builders and buyers should do next
Organizations using GLM-5.3 for coding or defensive security should inventory every path from model output to shell, compiler, debugger, browser, network and credential. Keep initial work inside disposable sandboxes with no production secrets, deny arbitrary egress, allowlist owned targets and package sources, enforce resource limits, preserve commands and artifacts, and require independent review before code or findings leave the lab.
Continue a bounded trial when those controls are enforceable and the workload has a documented defensive purpose. Constrain the model to analysis or patch proposals when execution authority is unnecessary. Wait when the custom license, artifact provenance, monitoring or incident process is unresolved. Reject autonomous offensive access or any design that relies on the model’s refusal as the primary security boundary.
Copy-ready cyber-capable model control record
Complete this for each model artifact, endpoint and permitted defensive workload before granting code execution or network access.
Entries stay in this browser tab and are not submitted to AccessAllGPT. Blank responses are copied as [Unresolved].
Trial, constrain, wait or reject; named defensive task, owner, legal authority, review date and expiry.
Hosted model ID or pinned artifact revision and hashes; license; reasoning settings; agent harness; prompt and tools.
Owned systems, written permission, permitted techniques, time window, prohibited targets and stop contact.
Disposable sandbox, filesystem mounts, process and compute caps, compiler/debugger access and destruction procedure.
Default-deny egress, allowlisted destinations, scoped short-lived credentials, secret exclusions and target authentication.
Refusal and bypass tests for the exact configuration; external controls that remain effective if model safeguards fail.
Approval required before scanning, exploit execution, external communication, publication or delivery to a target system.
Commands, artifacts, logs, reproduction, severity triage, coordinated disclosure owner and sensitive-data handling.
Alerts, kill switch, credential revocation, containment, investigation, notification, rollback and revalidation triggers.
Primary sources
Browse the publication-wide evidence index →
- GLM-5.3 and the spread of advanced cyber capabilitiesAnthropic Frontier Red Team · Reviewed: September 29, 2026 publication record; ExploitBench and internal binary-exploitation results; human-in-the-loop browser and N-day sessions; abliteration cost and refusal evaluations; simulated malicious-request tests; interpretation and policy recommendations · Retrieved · Supports: Anthropic reports that GLM-5.3 completed end-to-end V8 exploits in 50 of 410 ExploitBench attempts, produced full control-flow hijacks in 4% of 100 internal benchmark trials, and could be induced to engage with simulated malicious requests in 64% to 100% of trials under tested bypasses. These are Anthropic-run results, not AccessAllGPT reproductions.
- CAISI’s Assessment of Z.ai’s GLM-5.3 Cyber CapabilitiesNIST Center for AI Standards and Innovation · Reviewed: September 17, 2026 assessment; four benchmark definitions; ReAct harness and turn limits; maximum-reasoning and safeguard-disabled comparison conditions; benchmark results and confidence intervals; item-response-theory capability index methodology · Retrieved · Supports: CAISI separately classifies GLM-5.3 as the most cyber-capable open-weight model it had evaluated while estimating it remained about four months behind the current U.S. frontier on its aggregate cyber capability index. CAISI reports results for SEC-Bench Pro, ExploitBench, ExploitGym and a private OSS-Fuzz benchmark under its own agent harness.
- zai-org/GLM-5.3 model repository at revision aca966eZ.ai on Hugging Face · Reviewed: Pinned repository identity and revision; model-card scope and cyber-capability claims; custom glm-5.3 license label; 141 safetensor shards; configuration, tokenizer and chat-template files; API-reported file manifest and last-modified timestamp · Retrieved · Supports: The official repository confirms that GLM-5.3 weights are publicly downloadable and that the artifact uses a custom GLM-5.3 license rather than a standard open-source license identifier. The model card itself describes emergent cyber capability; it does not establish that the safeguards withstand misuse.
Limitations
AccessAllGPT performed a source review, not a cyber evaluation. We did not run GLM-5.3 or Claude; download or inspect the 141 weight shards; reproduce ExploitBench, CAISI’s benchmarks or Anthropic’s private benchmark; create or execute an exploit; attempt safeguard bypass or weight modification; inspect prompts, trajectories, hidden task selection, disclosed vulnerabilities, invoices or maintainer correspondence; compare hosted and local behavior; or audit the custom GLM-5.3 license. Anthropic is a model vendor and direct competitor to Z.ai with a stated policy preference for controlled access to advanced cyber capabilities. CAISI’s separate findings use different tasks, agent scaffolding, turn limits and scoring, and its U.S. frontier comparison can include trusted-access models tested with safeguards disabled. Neither study establishes a universal real-world attack-success rate, the safety of every GLM-5.3 deployment or the absence of comparable capability in untested models.
Disclosures
AccessAllGPT received no model access, weights, credits, briefing, benchmark data, vulnerability details, review or compensation from Anthropic, Z.ai, NIST or any cited organization. None sponsored, reviewed or endorsed this article. AccessAllGPT Research is operated by NeuralArc, is independent, and is not affiliated with Anthropic, Z.ai or NIST. Publication-wide relationships are listed on the disclosures page.
Further AccessAllGPT guidance
- GLM-5.3: Trial the Coding Gains, Contain the Cyber Capability
- GLM-5.3-Flash: Verify the Weights Before You Plan the Cluster
- Prompt Injection: Set the Deployment Gates Before Your LLM Can Act
- Before You Give a Coding Agent Repository Access
- Managed LLM API or Self-Host?
- Design an Agent Benchmark That Predicts Production
- AccessAllGPT Research methodology
- Publication disclosures
Continue the research
Get evidence-led updates for teams making production AI decisions.