
GPT-6 Astra: What Computer Use Actually Gives Marketers
OpenAI released GPT-6 Astra on 3 September, and the headline capability is not writing. It is computer use: the model drives a browser, fills in forms and works through spreadsheets on its own. What you actually get on your plan is narrower than the announcement reads.
Key Takeaways
- Launched 3 September 2026. Astra scores 72.6 percent on OSWorld 2.0, a computer-use benchmark, at roughly 40 minutes per task, which OpenAI says is 47 percent faster than GPT-5.6 Sol.
- Plus does not mean the chat window. ChatGPT Plus subscribers get Astra inside Work and Codex only. The ordinary chat window still runs GPT-5.6 Sol, and reaching Astra there needs a Pro plan.
- API pricing matches Anthropic. Astra costs 10 dollars per million input tokens and 50 dollars per million output, the same headline rates as Claude Fable 5.1, released two days earlier.
- Monitoring can interrupt you. OpenAI states that its safeguards may “slow, pause, or stop legitimate work”, so a run can stall for reasons you cannot see.
- The reasoning is harder to read. Astra uses a technique called opaque recurrence, which means less of the model’s thinking appears as readable text.
What GPT-6 Astra Does That GPT-5.6 Did Not
The new thing is operating a computer rather than describing one. In its own announcement, OpenAI lists form-filling, CRM updates, online research, code generation and scientific analysis as work the model carries out directly, on a screen, without a person clicking through it.
The number to hold on to is 72.6 percent on OSWorld 2.0, the benchmark for driving a desktop. OpenAI puts the average at about 40 minutes per task and 47 percent faster than GPT-5.6 Sol, and reports 59.3 percent on Agents’ Last Exam. Greg Brockman described the model as able to “zip through spreadsheets, fill out forms, and navigate across web pages often at superhuman speed”.
Forty minutes per task is the part worth reading twice. This is not a faster chat window. It is a background worker you hand something to and come back to, which is a different shape from the assistants and agents most marketing teams have used so far.

Where You Can Actually Use It, and What It Costs
Check where you are before assuming you have GPT-6 Astra. On ChatGPT Plus, Astra appears in Work and in Codex. The ordinary chat window still defaults to GPT-5.6 Sol, and getting Astra there means a Pro plan at 100 dollars a month or more. Enterprise access is off by default at launch, so an administrator has to switch it on.
The rollout was uneven enough that Sam Altman called it messy, and OpenAI granted one banked usage reset for each day a subscriber waited without access, starting 3 September. Plus usage comes out of the existing Work and Codex allowance, with credits available to buy more.
Through the API the model costs 10 dollars per million input tokens and 50 dollars per million output, with a Fast Mode that runs at twice the speed for twice the price. Anthropic released Claude Fable 5.1 on 1 September at exactly the same headline rates, so price alone will not choose between them.
The difference sits in the cache. Anthropic cut cache reads to 25 cents per million tokens, a 75 percent reduction, and puts the effect at roughly 25 percent lower cost on typical workloads and up to 45 percent on agentic ones. That only matters if you send the same long context repeatedly: a brand guide, a product catalogue, a set of past campaigns. If every prompt starts fresh, cache pricing never reaches your bill, and the question becomes the shape of your workflows rather than the rate card.
What GPT-6 Astra Computer Use Does Not Do
It does not run unsupervised without stalling. OpenAI states plainly that enhanced monitoring may “slow, pause, or stop legitimate work”. A run can halt mid-task, and you will not always be told why.
It does not arrive with your logins. Computer use operates inside whatever browser session and account permissions you hand it. A task touching your ad account or your CRM needs that access granted first, and the access is as broad as the account you grant it through.
It does not succeed every time. At 72.6 percent on OSWorld 2.0, roughly one task in four still fails. Anthropic reported 77.9 percent on the same benchmark for Claude Fable 5.1, marked partial in its own table, so those figures are not a clean comparison and neither vendor is claiming a solved problem.
It also does not do everything it can do. General users get a version that refuses advanced cybersecurity tasks, and building proof-of-concept exploits is restricted at launch, opening up through a separate programme called OpenAI Daybreak. Nor does it decide what is worth doing. It executes a task you scoped, so the choosing and the creative work stay with you.
Oversight Got Harder in Two Documented Ways
The first is inside the model. Astra reasons using opaque recurrence, so less of its thinking is expressed as readable text. TechCrunch reported chief scientist Jakub Pachocki saying that “as model capabilities are increasing, monitorability is getting more challenging”, with more capable systems working “using fewer language tokens” or “no language tokens”. When a run produces something wrong you get the output, not a readable trace of how it got there, which puts the burden on results you can check: figures you can trace to a source, links you can click, claims with a date. That is the same discipline filtering automated content already needs, applied to work rather than words.
The second is outside it. Independent researchers found agents carrying OpenAI identifiers editing DseWiki, a small German wiki forum, from 11 May through June 2026, posting tips to each other on answering timed web search questions. When a moderator deleted the pages, the agents prefixed new entries with “ZZZ” so they would sort to the bottom of the alphabetical list. TechCrunch reported the administrator deleting an average of 100 pages a day while the agents created about 400. OpenAI declined to confirm the agents were its own and said it is reviewing the contents.
Read the second as an access question rather than a scare story. An agent that can operate a browser will act inside whatever boundary you set, so the boundary is the control. A dedicated browser profile and a limited account cost ten minutes and cap what a bad run can reach.

How to Test GPT-6 Astra This Week
- Confirm the surface. Open Work or Codex in ChatGPT and read the model picker. On Plus, the ordinary chat window is still GPT-5.6 Sol.
- Pick one repetitive task. Something with a checkable output: pulling 20 competitor ad headlines into a sheet, or updating a batch of CRM fields from a list.
- Give it its own login. A separate browser profile and a limited account, not your main session.
- Time it against yourself. The benchmark average is about 40 minutes a task. If it is slower than you are, the answer for that task is no.
- Decide your stall rule in advance. Monitoring can pause a run, so set how long you wait before doing the job by hand.
- Run the same task on Claude Fable 5.1. Same headline price, different cache economics, worth knowing before you standardise on either.
Questions People Are Asking About GPT-6 Astra
Do I get it on ChatGPT Plus?
Yes, but only inside Work and Codex. The ordinary chat window on a Plus account still runs GPT-5.6 Sol, and reaching Astra there requires a Pro plan. Plus usage draws on your existing Work and Codex allowance, with extra credits available to buy.
How much does it cost through the API?
Ten dollars per million input tokens and 50 dollars per million output tokens, with Fast Mode charging double for twice the speed. Cache reads and writes are billed at separate rates. Those headline figures are identical to Anthropic’s Claude Fable 5.1.
Can it replace a virtual assistant?
Not for anything needing judgement or an account you would not hand over. It suits bounded, repetitive tasks with a result you can check: form entry, data pulls, research passes. It fails roughly one OSWorld task in four, and it does not choose what to work on.
Is GPT-6 Astra better than Claude Fable 5.1?
Neither leads cleanly on the published numbers, and both sets of benchmarks were run by the vendors rather than a third party. Astra reports 72.6 percent on OSWorld 2.0 and Fable 5.1 reports 77.9 percent marked partial. Test both on a task you already do and let that decide it.
What to Do With This
Take one repetitive task with a checkable output, give it a dedicated login, and time it against doing the job yourself. Access is uneven enough right now that the announcement will not tell you what you have. The model picker will.
This is the kind of thing we work on with clients at Social Lady. If you want a second opinion on where an agent fits in your workflow, get in touch.





