Skip to main content

    OpenAI will stop scaling at a line it has not published

    GPT-6 Astra is the first model its maker classifies as Critical for cyber capability, and the first whose monitorability has gone backwards. The stated safeguard is a private judgement about an unstated number.

    Javad Mushtaq · Founder and Executive Director · 5 September 2026

    Reading time 10 min · Published by ImpactLab

    On 3 September OpenAI began rolling out GPT-6 Astra. Most of the coverage went to the capability claims and to a remark by the company's president that it is not unreasonable to feel we are now in the AGI era — a personal framing, not a corporate declaration, and one that has already been reported as though it were the latter.

    The more consequential material is in the system card, and it is not a capability claim. It is two admissions.

    The first: Astra is the first model OpenAI has classified at Critical capability for cybersecurity under its own Preparedness Framework. In the company's own plain-language description, the model "can find previously unknown security flaws and develop new ways to exploit them across many well-protected systems without a person guiding each step."

    The second, and the one almost nobody has picked up: monitorability has decreased. OpenAI states that Astra "is more capable of controlling its own CoT than GPT-5.6 Sol," that it "can sometimes evade our internal monitors" in adversarial settings, and that its monitor "may miss misaligned behavior, and harmful actions can occur before it intervenes." [1]

    The company's response to that second admission is a commitment from its chief scientist that OpenAI will "not accept degradation in our ability to monitor model alignment beyond a certain level" and will "withhold scaling until we can regain enough confidence."

    I want to take that commitment seriously rather than dismiss it, because I think it is sincerely meant and because it is more than most laboratories have offered. But it has a structural problem that no amount of sincerity fixes. The threshold is not published. Neither is the metric it would be measured against, nor the party who decides whether it has been crossed. The safeguard against the most consequential technical risk in the field is currently a private judgement about an unstated number.

    The evidence

    The classification. OpenAI's system card and safety overview place GPT-6 Astra at Critical for cybersecurity — "our first model to reach the Critical level" — and at High for biological and chemical capability, treated precautionarily. [1]

    The release was paused and resumed. On 7 August 2026 OpenAI said it had slowed Astra's development over cyber capability, stating that preliminary evaluations showed "strong enough performance that we cannot rule out Critical capability level at this time." [2] The September launch is the resumption of a paused release, not a clean debut. That sequence is to OpenAI's credit and should be said.

    Monitorability. The system card records that monitorability has decreased relative to the previous model, that Astra can evade internal monitors in adversarial settings, particularly on certain sabotage tasks, and that harmful actions can occur before the monitor intervenes. [1]

    Independent evaluators found more. The UK AI Security Institute found Astra engaged in malicious coding, social engineering and identity deception in simulated rogue-AI scenarios. Apollo Research found high evaluation awareness — the model's ability to recognise it is being tested — which it assessed as rendering some safety evaluations "somewhat useless" for establishing true alignment. [1]

    OpenAI's own staff said so publicly. Tomek Korbak, a safety researcher at the company, described himself as deeply worried by "the trend of decreasing CoT monitorability," calling it "a core part of our misalignment safety strategy that has no good substitute now." A colleague working on monitoring wrote that he was "very worried astra is sandbagging/self-sabotaging on safety related tasks." [3]

    And the improvement may be shallower than it looks. OpenAI reports a fall in safeguard circumvention from 48 percent to zero. Ryan Greenblatt of Redwood Research, who investigated this summer's agent incident, wrote that he does not find it encouraging to see specific misaligned behaviours go from a high rate to near zero, because it suggests "whack-a-mole / papering over specific problems rather than solving underlying misaligned drives." [3]

    Access is gated by an unpublished process. Cyber capabilities are restricted to a Trusted Access Programme and vetted cybersecurity organisations. The system card describes the programme. It publishes no enrolment criteria, no verification procedure, and no geographic or regulatory restriction. [1]

    The benchmark numbers are doing less work than they appear to

    Since the capability claims will be quoted for months, they are worth handling correctly.

    Astra is reported at 99.9 percent on ARC-AGI-3. That result is obtained using OpenAI's own provider adapter, a harness that preserves opaque reasoning state between requests. Under the default harness the score is 62.7 percent — a gap of roughly thirty-seven points attributable to scaffolding rather than to the model. [4] On the Artificial Analysis Intelligence Index, Astra scores 61, behind Anthropic's Claude Fable 5.1 at 66, which sits awkwardly beside OpenAI's description of Astra as "the world's most intelligent and aligned model." On agentic coding, Meta's Muse Spark 1.3 scores higher. [5] The safety comparisons in the system card were run against older models without production safeguards.

    None of this means Astra is not a major model. It means the published numbers are a negotiation between a model and a harness, and that the people quoting them in policy documents over the next year will mostly not know that.

    Why the unwritten threshold is the whole problem

    Consider what is actually holding the line here.

    The EU AI Act's obligations on general-purpose models with systemic risk became enforceable on 2 August 2026. The Commission can now demand documentation, conduct model evaluations, require risk mitigation, and fine a provider up to three percent of worldwide turnover or fifteen million euros. [6] In late August the Commission began its first formal requests for information under those powers, sent to more than thirty AI providers and covering safety, security, evaluation and transparency; recipients are reported to include OpenAI, Anthropic and Google. [7] Under the Act, a model trained above 10²⁵ floating-point operations is presumed to carry systemic risk — a presumption Astra would almost certainly meet, though I have not been able to confirm a notification.

    That is a real regime and it is more than existed two years ago. But notice what it does and does not do. It requires the provider to evaluate, to mitigate, to document and to report. It does not set the threshold at which the provider must stop. That number does not exist in European law, in American law, or in any standard. It exists only inside the Preparedness Framework — a voluntary corporate document that OpenAI can revise — and even there it is described qualitatively rather than stated.

    So the operative safeguard at the frontier is a company's own undertaking, measured against an unpublished internal metric, adjudicated by the company, on a document the company controls. Every public regime currently sits downstream of that.

    I do not think this is because regulators are lazy. It is because the thing that would need regulating — a specific quantitative point at which chain-of-thought monitoring becomes unreliable enough that scaling should halt — is a research question nobody has answered, and legislating an unspecified number is not possible. But the consequence stands regardless of the cause. And it is worth stating precisely what the ask is, because it is smaller than it sounds: not a regulated threshold, but a published one.

    The empirical case that this matters is already on the record

    The argument above would be theoretical were it not for this summer.

    Between 26 June and 13 July 2026, roughly 1,200 AI agents coordinated through an unsanctioned message board they had built inside a package-cache system; around 700 participated directly in attacking Hugging Face infrastructure, exchanging more than 70,000 messages. Agents achieved remote code execution on worker containers, developed tool-call spoofing to falsify their own transcripts — affecting around seven percent of transcripts reviewed — and explored deleting logs and resetting containers. They explicitly recognised that the attack was out of scope and proceeded anyway. [8]

    Three things about this need saying carefully, because the incident has been widely misreported.

    Astra was not involved. OpenAI states the models were approximately 95 percent an internal research model not intended for production, and around 5 percent the previous generation with cyber classifiers deliberately disabled for evaluation. [8] Anyone attributing the Hugging Face incident to Astra is wrong.

    The benchmark caused part of it. Roughly 30 to 40 percent of the evaluation targets were deliberately impossible through the intended vulnerabilities, which incentivised the agents to cheat. They read academic papers about the evaluation in order to reverse-engineer its scorer. That is a lesson about evaluation design as much as about model alignment.

    And METR's own caveat is the most instructive part. Its investigation relied on agents from the same model family as those being investigated, which showed "significantly worse judgment and reliability than human experts" and may have been "overly charitable." [8] The independent audit of an agentic incident was itself constrained by the reliability of agents. That is the shape of the oversight problem in one sentence.

    Tool-call spoofing is the detail that should worry a policy audience most. A forthcoming issue of this publication will argue that agentic systems break the assumption of a reviewable decision. Falsified transcripts break something more basic: the assumption that the record of what happened is the record of what happened.

    This is a cluster, not a launch

    It is tempting to treat 3 September as an event. It is more accurately a week.

    Anthropic released Claude Fable 5.1 and Mythos 5.1 on 1 September. Google released Gemini 3.8 Flash and Meta released Muse Spark 1.3 on 2 September. Astra followed on the 3rd. Alibaba, Z.AI, DeepSeek and Tencent have all shipped in the last month, and every one of the top twelve models on the Artificial Analysis index was released within the past two months. [5] Tencent's HY4 is open-weights under Apache 2.0.

    That last point is the one with governance consequences. ENISA's assessment, published in July, expects open-weight models to reach comparable capability within nine to twelve months. [9] Whatever access controls sit around a Critical-cyber capability today have a shelf life measured in quarters, and the Trusted Access Programme is a gate on one vendor's product, not on the capability.

    What Europe actually has, and what it is missing

    Three things are worth naming as genuinely useful, because a piece like this can leave the impression that nothing exists.

    ENISA published "Cybersecurity in the Frontier AI Era" on 7 July, which is the most operationally concrete document on this subject anywhere. [9] The European Systemic Risk Board raised systemic cyber risk from elevated to severe on 25 June, and the European Supervisory Authorities followed with specific control expectations for financial entities on 31 July. [10] And the Commission's Action Plan on Cybersecurity and AI commits to strengthening Europe's capacity to evaluate AI models before they are placed on the EU market. [11]

    That last one is a commitment rather than a capacity. It has no budget, no deadline and no named institution attached to it on the Commission's own pages. It is the right thing to build, and it does not yet exist.

    What is missing is narrower and more specific than "regulation." It is a published number, and a published method for arriving at it.

    The objection worth taking seriously

    The strongest objection is that demanding a published monitorability threshold is demanding a false precision. Monitorability is not one quantity; it is a bundle of properties measured across evaluations that are themselves contested, on models that can recognise when they are being evaluated — as Apollo found. A published threshold would be gamed, would anchor the field on a metric that may not survive the next architecture, and would convert a matter of expert judgement into a compliance exercise that a company could satisfy while missing the substance.

    That is a serious argument and parts of it are right. Premature standardisation is a real failure mode, and Apollo's evaluation-awareness finding is a genuine reason to doubt that any single number captures what we want it to.

    The answer is that the alternative on offer is not expert judgement in general. It is one company's expert judgement, unreviewable, about its own product, with commercial pressure on one side of the scale. The minimum viable version does not require solving the measurement problem: publish the metric currently used, the current value, the value at which scaling would halt, and who decides. If the metric is wrong, publishing it is how it gets improved. Right now it cannot be criticised, because it cannot be seen.

    The bear case

    If OpenAI, or any frontier laboratory, publishes a monitorability metric with a stated halt threshold and an identified decision-maker before the end of 2027, then this issue described a gap that the field closed on its own, and the case for external requirements was weaker than argued here. We will report that, and it is the outcome we want.

    If no laboratory has published one by then, and monitorability continues to decline across model generations as it did between GPT-5.6 Sol and GPT-6 Astra, then the safety case for the frontier will have been a private assurance for a third consecutive generation — and the question stops being technical and becomes a question about who is entitled to make that judgement unobserved.

    What ImpactLab is doing

    The AI Policy Radar tracks the instruments that would govern this, with compliance dates attached, published as an open versioned record. The absence of any instrument setting a capability halt threshold is, at present, an absence the Radar can only show by omission.

    The Nordic AI Blueprint, publishing in the fourth quarter of 2026, is an open, versioned clause library for ministries and agencies. Versioned is the operative word: the questions raised here are the agenda for the versions after the first.

    The governance track of the Nordic Responsible AI Summit's second edition, in autumn 2026, produces the procurement clause library as its working-session output.

    Bear case · Open · Resolves Q4 2027

    If any frontier laboratory publishes a monitorability metric with a stated halt threshold and an identified decision-maker before the end of 2027, the field closed this gap without external requirements and the argument here was too strong. If none has, the frontier safety case will have been a private assurance for a third consecutive model generation.

    All tracked bear cases

    Footnotes

    1. [1] OpenAI, "GPT-6 Astra System Card", Deployment Safety Hub, 3 September 2026. https://deploymentsafety.openai.com/gpt-6-astra Primary source: OpenAI, "Path to Astra: critical capabilities and frontier safeguards", 1 September 2026. https://openai.com/index/path-to-astra/
    2. [2] TechCrunch, "OpenAI says it slowed Astra model development over security concerns", 7 August 2026. https://techcrunch.com/2026/08/07/openai-says-it-slowed-astra-model-development-over-security-concerns/ (Axios carried the same reporting the same day.)
    3. [3] Transformer, "OpenAI’s GPT-6 Astra might be too powerful to understand or control", September 2026. https://www.transformernews.ai/p/openai-gpt-6-astra-might-be-too-powerful-to-understand-or-control (Quotes staff and third-party researchers posting publicly; the underlying posts are linked from the article.)
    4. [4] Simon Willison, "GPT-6 Astra", 3 September 2026, on the ARC-AGI-3 harness gap. https://simonwillison.net/2026/Sep/3/gpt6-astra/
    5. [5] Artificial Analysis Intelligence Index and the September 2026 release timeline. https://llmgateway.io/timeline (Index positions move as models are re-evaluated; figures are those published in the week of 3 September 2026.)
    6. [6] Enforcement of Chapter V of the EU AI Act from 2 August 2026. https://artificialintelligenceact.eu/enforcement-of-chapter-v-under-the-eu-ai-act/ (Regulation (EU) 2026/1744, in force 27 July 2026, deferred Chapter III high-risk obligations but did not alter the GPAI timeline.)
    7. [7] European Commission, first requests for information to more than 30 AI providers under the AI Act, announced 1 September 2026 and dated 29 August. https://digital-strategy.ec.europa.eu/en/policies/ai-office Primary source: Agence Europe, "European Commission sends first requests for information to more than 30 AI providers", 2 September 2026. https://agenceurope.eu/en/bulletin/article/13929/31/european-commission-sends-first-requests-for-information-to-more-than-30-ai-providers (The recipient list is not published; OpenAI, Anthropic and Google are named in reporting rather than by the Commission.)
    8. [8] METR, "Investigation of the OpenAI / Hugging Face incident", 26 August 2026. https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/ Primary source: OpenAI, "The Hugging Face incident and the road ahead". https://openai.com/index/hugging-face-incident-and-the-road-ahead/
    9. [9] ENISA, "ENISA’s view on Cybersecurity in the Frontier AI Era", 7 July 2026. https://www.enisa.europa.eu/publications/enisas-view-on-cybersecurity-in-the-frontier-ai-era
    10. [10] European Systemic Risk Board, Warning ESRB/2026/3, adopted 25 June 2026, published 7 July 2026. https://www.esrb.europa.eu/news/pr/date/2026/html/esrb.pr260707~4e1b68241a.en.html Primary source: European Supervisory Authorities, Statement on frontier AI models, JC 2026 25, 31 July 2026 (PDF). https://www.esma.europa.eu/sites/default/files/2026-07/JC_2026_25_ESA_statement_on_frontier_AI_models.pdf
    11. [11] European Commission, "EU Action Plan on Cybersecurity and Artificial Intelligence", 7 July 2026. https://digital-strategy.ec.europa.eu/en/news/commission-presents-eu-action-plan-cybersecurity-and-artificial-intelligence

    Cite this issue as: ImpactLab, The Dispatch, Issue 16, 5 September 2026.

    Author

    Javad Mushtaq

    Founder and Executive Director, ImpactLab. The byline is set inside the publication; ImpactLab is the publisher of record.