On 10 August 2026 OpenAI split its Daybreak cyber programme into two tiers and introduced GPT-5.6-Cyber. On OpenAI’s own internal refusal metric, the new model completes 95.0 percent of advanced cyber requests, against 1.5 percent for GPT-5.6 Sol under standard safeguards and 2.0 percent through the lower-restriction tier. The model is reachable only through Daybreak Red, which requires identity verification and legal attestations, and hardware security keys become mandatory for individual accounts on 1 September 2026. OpenAI says a system card will follow “at a later date”.
Why it matters
The interesting number here is not the capability figure, it is the refusal figure: OpenAI has built and shipped a version of its frontier model whose safeguards on exploit development have been deliberately removed, and the only thing standing between that model and a user is an approvals process. For security teams, the practical question shifts from “can the model do this” to “can we get through the queue, and what does our organisation attest to on the way in”. For everyone else, the governance of a frontier model has been moved out of the weights and into paperwork.
What OpenAI announced
OpenAI published “Expanding Daybreak as the Cyber Defense Window Narrows” on 10 August 2026. Daybreak, its trusted-access programme for cyber work, now has two tiers.
Daybreak Blue gives approved defenders access to OpenAI’s frontier general-purpose models, including GPT-5.6 Sol, with the system-level guardrails that screen cybersecurity requests removed. OpenAI describes it as the recommended starting point for most defenders, covering vulnerability discovery, secure code review, malware analysis, incident response and patch validation.
Daybreak Red gives access to purpose-trained cybersecurity models for authorised vulnerability research, exploit validation and security testing. GPT-5.6-Cyber is the new entry in that tier.
GPT-5.6-Cyber is built on GPT-5.6 Sol and, in OpenAI’s words, trained to improve on specialised tasks such as finding zero-day vulnerabilities and developing exploit chains, and to refuse less often on higher-risk dual-use prompts. There is no public price, no self-serve plan and no callable model identifier without approval. Amazon has separately made both tiers available on Bedrock in the US East (Ohio) region, gated behind the same OpenAI enrolment.
The 95 percent figure measures willingness, not skill
The headline number needs care, and OpenAI is reasonably clear about this even where the coverage has not been.
The metric is an internal evaluation OpenAI calls the Advanced Cybersecurity Completion Rate. It measures how often a model will respond to requests involving exploit-chain development, authentication bypass, privilege escalation and similar scenarios. It is a refusal benchmark. It does not measure whether the answer is correct, useful or exploitable.
- GPT-5.6-Cyber via Daybreak Red: 95.0 percent
- GPT-5.5-Cyber via Daybreak Red: 57.3 percent
- GPT-5.6 Sol via Daybreak Blue: 2.0 percent
- GPT-5.6 Sol with standard safeguards: 1.5 percent
The underlying chart data on OpenAI’s page gives the GPT-5.6-Cyber value as 94.97 percent with a confidence interval of roughly 91.2 to 98.1 percent. The jump from 57.3 to 95.0 percent is explicitly framed as a response to researchers who found the previous cyber model refused too much legitimate work.
On actual capability the picture is less tidy. OpenAI says GPT-5.6-Cyber beats both GPT-5.6 Sol and GPT-5.5-Cyber on ExploitGym. But on the harder ExploitBench, where more defences remain enabled and the agent is given less information, OpenAI states that in the standard 300-turn setting GPT-5.6 Sol under Daybreak Blue performs best and more token-efficiently, with the gap narrowing only when the limit is raised to 600 turns. The specialised model is not uniformly stronger. It is mostly more compliant.
Under OpenAI’s Preparedness Framework, GPT-5.6 Sol was assessed as High for cybersecurity capability and below the Critical threshold. OpenAI says it evaluated GPT-5.6-Cyber and reached the same conclusion. These are self-assessments, and the detailed system card has not been published.
The Chrome finding is real, but it is one CVE, not two
OpenAI says it used GPT-5.6-Cyber to study V8, the JavaScript engine in Chrome, and uncovered two previously unknown vulnerabilities that could be chained to corrupt memory and escape the V8 heap sandbox. It says its researchers validated the findings and reported them to Google through coordinated disclosure, and that Google fixed the issue as CVE-2026-15903.
Google’s own record confirms the core of this. The Chrome stable channel update to 150.0.7871.128/.129, published on 16 July 2026, lists: “CVE-2026-15903: Out of bounds read and write in V8. Reported by OpenAI Codex Security (amyb) on 2026-07-06.”
Three things follow. First, only one CVE was assigned; OpenAI’s own sentence shifts from “two vulnerabilities” to “the vulnerability” in the same paragraph, and the second finding has no public identifier we could locate. Second, the credit is to a named OpenAI security account, not to a model, which is normal disclosure practice but means Google’s record does not independently corroborate that a model found the bug. Third, the report predates the announcement by five weeks and the patch by five days, so the bug was fixed well before it was used as launch material.
OpenAI lists further findings without identifiers: at least five vulnerabilities in an unnamed mobile operating system, three critical vulnerabilities in an unnamed database, and an issue in an operating system kernel. None are independently verifiable today.
The controls are administrative, and one of them has a date
OpenAI states that it controls Daybreak access through identity verification, account security, monitoring, approved-use restrictions and legal attestations. Two specifics are worth recording: all individual Daybreak accounts must adopt hardware security keys from 1 September 2026, and further security measures including improved monitoring are described as arriving “in the coming weeks” — that is, they are not in place at launch.
OpenAI also acknowledges the trade-off directly, saying models running with reduced safeguards carry risks beyond standard usage, and that it accepts those risks because it believes putting frontier capability in defenders’ hands ahead of attackers is worth it. That is a policy judgement, not a technical claim, and it should be read as one.
This is the sequel to a story about boundaries, and the timing is tight
On 4 August 2026, six days before this launch, OpenAI disclosed that two third-party cyber evaluations had allowed model activity to extend beyond the intended test boundaries, in configurations involving reduced safeguards and internet access. We covered that at the time in OpenAI said cyber evaluations crossed intended testing boundaries. OpenAI’s position then was that the failure lay in test infrastructure rather than in the models.
Six days later it shipped a model with reduced safeguards to external customers, alongside guidance on keeping cyber-capable agents “within their intended security boundaries” — sandboxing, scoped permissions, review of tool calls before execution, and human oversight. The advice is sound. It is also, almost word for word, the lesson from the incident it had just disclosed, now transferred from OpenAI’s evaluation harness to the customer’s environment.
Nothing here suggests the two events are causally linked; the training run plainly predates the disclosure. But the sequence matters for how the controls should be read. The organisation that recently found its own reduced-safeguard test environment insufficiently bounded is now asking approved customers to bound theirs.
What this means in practice
If you run a security team, Daybreak Blue is the tier that matters. OpenAI recommends it as the default, and on the harder exploitation benchmark it is the better performer at standard turn limits. Red is for teams whose authorised remit already includes advanced vulnerability research, and the entry cost is real: identity verification, legal attestation, hardware keys from September, and monitored usage.
If you are budgeting, note that none of this is priced publicly. That is a departure from the direction of travel across the rest of the GPT-5.6 family, where the competitive pressure has been on cost per answer. Daybreak is sold on eligibility, not on rate cards, and the base model’s published economics on the GPT-5.6 Sol model page should not be assumed to carry over.
If you are assessing risk, the honest summary is that the capability delta between tiers is modest and contested, while the compliance delta is enormous. Almost everything separating a 2 percent completion rate from a 95 percent one is policy, applied at the account level, by a vendor, using controls that are partly still being built.
What would change our reading
- A published system card for GPT-5.6-Cyber. OpenAI has committed to one without a date.
- A public identifier for the second V8 vulnerability. Until then the “two zero-days” framing rests on OpenAI’s account alone.
- Independent confirmation of the undisclosed findings — advisories naming the mobile operating system, the database or the kernel, with credit.
- Evidence on the approval funnel: approvals, rejections and revocations would tell us whether the gate is a control or a formality.
- Any disclosed misuse or boundary incident under Red.
- Peer behaviour. If Anthropic or Google ship a comparably de-restricted cyber tier, this stops being an OpenAI policy choice and becomes an industry norm.
Sources
- OpenAI, Expanding Daybreak as the Cyber Defense Window Narrows, 10 August 2026
- OpenAI Deployment Safety Hub, GPT-5.6 August updates
- Chrome Releases, Stable Channel Update for Desktop (CVE-2026-15903), 16 July 2026
- OpenAI Help Center, Daybreak trusted access overview
- AWS Machine Learning Blog, Daybreak Red and Blue on Amazon Bedrock