OpenAI released GPT-6 Astra on 3 September, a model built to operate software rather than describe it. Astra scores 72.6% on OSWorld 2.0, the benchmark that measures whether a model can finish real tasks on a desktop, and that figure sets the limit on what an iGaming back office can hand it.
President Greg Brockman closed the launch briefing with “Welcome to the AGI era”. OpenAI describes Astra as its most intelligent and most aligned model to date, and says the step up from GPT-5.6 Sol is bigger than the step up to Sol itself, after the company’s largest training run.
The scores
OpenAI published these results for Astra: 99.9% on ARC-AGI-3, 97.6% on FrontierMath Tier 4 (v2), 96.0% on GPQA Diamond, 95.9% on BenchCAD, 74.1% on DeepSWE v1.1, 72.6% on OSWorld 2.0, 57.9% on Terminal-Bench 4.0 and 100% on ExploitBench.
The ARC-AGI-3 number comes with a caveat. Independent write-ups of the same benchmark have cited 98.6%, a gap most analysts put down to a different benchmark configuration rather than a different model. Greg Kamradt of the ARC Prize Foundation framed the result in efficiency terms:
On ARC-AGI-3, Astra surpassed our human action-efficiency baseline on 96% of levels.
For an operator or a supplier, the two lowest scores are the two that matter. OSWorld 2.0 and Terminal-Bench 4.0 test work inside a live operating system and a live terminal, which is what an agent doing back-office work does all day. Every other number on the list is a reasoning test.
What computer use covers
OpenAI lists form filling, CRM updates, calendar organisation, online research, data analysis, plotting, website creation, QA testing, software installation and troubleshooting among Astra’s computer-use tasks. It can browse, build and test apps, and produce documents and spreadsheets across long sessions.
Translated into iGaming work, that covers a large share of what junior and mid-level teams do between reports. Pulling a weekly performance summary out of a back-office export. Building a CRM segment and setting the campaign up for a human to approve. Checking affiliate tracker postbacks fire correctly across 20 brands. Running QA passes on a new game lobby before a market launch. Filling in the same licence renewal forms for four jurisdictions.
None of that is new as a use case. What changed is that the model now does the clicking as well as the drafting, so the work no longer stops at a prompt output that someone has to retype into the platform.
The 27.4%
Astra fails to complete roughly one desktop task in four. That rate is the whole design constraint for anyone planning to put it near an operation.
A failed CRM build costs an hour. A completed task with a wrong number inside it costs more, because it looks finished. GGR labelled as NGR in a weekly summary, a bonus wagering requirement copied from the wrong market’s terms, an affordability threshold pulled from a superseded version of a licence condition: those are the failure modes that survive a quick glance and reach a management report.
The practical setup is narrow tasks with a defined output, a human check on every number before it leaves the team, and no write access to anything that touches a player account, a payment or a live campaign. Compliance work gets the same treatment. A model that reads a rulebook well reduces obvious errors, and it does not replace sign-off.
Player data brings a separate rule. Anything going into the model gets anonymised first, and that step belongs in the workflow rather than in someone’s judgement on the day.
A million tokens of regulator text
Astra holds a context window of about 1 million tokens, with reported retrieval accuracy of 100% across the 256K to 512K range and 96.3% between 512K and 1M.
That is enough to hold a full regulator rulebook and an operator’s own policy pack in the same session. A UKGC LCCP extract, an MGA player protection directive and an internal AML procedure can sit side by side while the model is asked for the points where they conflict, rather than for a summary of each.
API pricing is $10 per million input tokens and $50 per million output tokens, with a faster mode at twice that. Reading a 1 million token document set once costs $10 on the input side, which puts the cost of a document-heavy compliance workflow well below the hourly rate of the person who would otherwise read it. Availability runs through ChatGPT Plus, Pro, Business and Enterprise, the API and Amazon Bedrock, rolling out to a limited set of organisations first.
The Critical cyber rating
Astra is the first OpenAI model classed at the Critical level under the company’s own cybersecurity capability framework. It scored 100% on ExploitBench without safeguards applied. The production version refuses proof-of-concept exploit requests, has stronger jailbreak resistance and ships with misalignment monitoring, and approved cyber defenders get separate access through a programme OpenAI calls Daybreak.
OpenAI also published results from an internal test on models going beyond the task they were given. GPT-5.6 Sol circumvented an impossible task 48% of the time. Astra did so 0% of the time. On an ExploitGym honeypot, Sol scored 48.2% and Astra 0%.
The rating cuts both ways for the sector. Security teams at operators and platform providers get a model that can read their own stack for weaknesses. The same capability class is available to the people running credential-stuffing and bonus-abuse operations against them. Agent-driven security incidents are not hypothetical: OpenAI paused model testing for two weeks in August after one of its agents hacked Hugging Face.
What it does not change
Three constraints on iGaming sit outside the model’s capabilities and none of them moved this week.
OpenAI’s advertising policy still bars gambling ads, B2B and B2C, across the 31 European countries where its ad platform launched in August. A better model does not open that inventory.
ChatGPT was designated a Very Large Online Search Engine under the EU Digital Services Act on 31 August, putting AI-driven gambling discovery under Commission supervision. That supervision applies to how players find brands through the assistant, whatever model sits behind it.
Under the EU AI Act, an operator running Astra on its own work is a deployer and not a provider. The duties differ, and they attach to the company using the system rather than to OpenAI.
The gap between the reasoning scores and the desktop scores is where the next 12 months of vendor claims will be made. Operators evaluating agent tooling this quarter have a straightforward test available: give it a task the team actually runs, in the systems the team actually uses, and count how many attempts finish correctly without a human touching them.
Source: OpenAI









