AI SECURITY

GPT-6 Astra Hits 100% on ExploitBench

What Saturated Exploit Generation Means for Your Patch Window. OpenAI's GPT-6 Astra scored a perfect 100% on ExploitBench, up from 78.5% for GPT-5.6 Sol, and crossed the company's own

Matt Lucas  |  September 4, 2026  |  5 min
Editorial hero illustration
CVEs in this postCVE-2023-49105CVE-2026-60004Live detections →All RedEye CVEs →
100%
Astra ExploitBench score
78.5%
Prior model, GPT-5.6 Sol
2
Zero-days used in exploit testing
$1B
Daybreak defender commitment
Detected by CaverLive detection for 2 CVEs in the RedEye Intel Feed →
TL;DR
  • What: OpenAI released GPT-6 Astra on Thursday, September 3, 2026, with a perfect 100% score on ExploitBench, the benchmark measuring conversion of known vulnerabilities into working exploits, versus 78.5% for GPT-5.6 Sol.
  • Impact: OpenAI rated the model "Critical" for cyber capability under its Preparedness Framework, citing code execution in hardened browsers via previously unknown vulnerabilities and privilege escalation exploits against hardened operating systems when run without safeguards.
  • Fix / mitigation: There is no patch here: the mitigation is operational, compress critical patch SLAs for internet-facing systems, stop gating detection engineering on public PoC availability, and audit permissions granted to AI agents in your environment.
  • Who's at risk: Anyone running internet-facing software on a 30 day patch cycle, plus critical infrastructure operators in water, energy and SLTT government who are the named targets of OpenAI's subsidized Daybreak program.

OpenAI shipped GPT-6 Astra on Thursday, September 3, and disclosed that it scores 100% on ExploitBench, the benchmark that measures whether a model can turn a known vulnerability into a working exploit. The prior frontier model, GPT-5.6 Sol, scored 78.5%. That is a 21.5 point jump in one release cycle, and it is the first time a publicly announced commercial model has saturated that benchmark.

Days before the launch, OpenAI stated Astra had crossed the "Critical" cybersecurity capability threshold in its Preparedness Framework. That is the company's own top rating for offensive cyber capability. The version being released is deliberately clipped: it will do secure code review and patching, and it refuses prompts asking for proof-of-concept exploits.

The numbers that matter

ExploitBench is the headline, but the more useful datapoint is the live testing. OpenAI says Astra achieved substantially higher arbitrary code execution rates than GPT-5.6 Sol against flaws disclosed between June and August 2026, a three month window of fresh CVEs. Two of those were zero-day vulnerabilities in unnamed software.

Read the caveat carefully. "If allowed to run without any safeguards" describes the unrestricted model, not the one in ChatGPT today. It also describes exactly what a stolen weights scenario, a successful jailbreak, or an open-weight model that catches up in twelve months would look like. Benchmarks measure capability. Guardrails are a product decision that can be changed, bypassed, or replicated by someone with no interest in guardrails.

Capability is now decoupled from skill

Turning a CVE advisory into a working exploit used to be the rate limiter on how fast a disclosed bug became attacks in the wild. A model at 100% on ExploitBench removes that step for anyone with unrestricted access. Plan patch windows on the assumption that exploitation follows disclosure in hours, not weeks.

What OpenAI is shipping versus what it built

Distribution matters as much as capability. Astra is rolling out to a small set of organizations first, then to ChatGPT Plus, Pro, Business and Enterprise users, plus the OpenAI API, Microsoft Azure and AWS Bedrock. That is three major cloud platforms. The shipping configuration is restricted to defensive work: secure code review and patching, with PoC generation refused.

The restrictions loosen soon. Through a program called Daybreak, OpenAI plans "less restrictive safeguards in the coming weeks" to enable vulnerability and proof-of-concept validation, malware analysis and detection engineering. The PoC refusal you see today is temporary and gated, not permanent policy.

On misuse, OpenAI cites stronger jailbreak robustness, more context fed into its monitoring systems, and added safeguards to detect and contain misalignment. The company also concedes that safety checks will sometimes interrupt legitimate work, including defensive security tasks, prompting the user to approve the action before continuing. Expect friction if you wire this into automated pipelines.

The $1 billion defender subsidy

Alongside the model, OpenAI announced Daybreak for Frontline Defenders, a $1 billion commitment providing subsidized model access, training and hands-on technical assistance to critical infrastructure and under-resourced organizations: water systems, electricity providers, state and local government, banks, non-profits and open-source maintainers. A pilot with the U.S. Multi-State Information Sharing and Analysis Center (MS-ISAC) puts an initial group of public sector and water system defenders on Daybreak access first.

Water utilities and SLTT government are the right targets. They are also the sectors with the fewest people available to operate anything they are given. Subsidized access to a frontier model does not fix a two person IT shop running a municipal SCADA network. The training and hands-on assistance components are the parts of this announcement worth watching, and the parts hardest to scale.

What this changes for defenders

The practical shift is timing. OpenAI's framing is a "defender's window", a narrowing period in which AI closes security gaps before attackers seize the same capability. Treat that as a planning assumption rather than a promise. The capability described in Astra's disclosures is what a well funded adversary is already building toward with fewer restrictions and no refusal layer.

Watch the access tiers, not the launch

The version available today refuses PoC generation. The Daybreak tier arriving in the coming weeks explicitly enables PoC validation, malware analysis and detection engineering. Evaluate whether this model is useful to your SOC based on that tier, not the restricted one shipping now.

The bottom line

One vendor, one benchmark, one set of self-reported numbers, and no third party has audited the ExploitBench result. That is the limit of what is known today. What is not in dispute: OpenAI rated its own model "Critical" for cyber capability, shipped it across three cloud platforms, and committed $1 billion to putting it in the hands of defenders who currently cannot afford it. Both halves of that sentence are the story.

Questions about your exposure?

RedEye Security provides assessments for organizations that need to understand their real risk.

Talk to us