OpenAI launches GPT-6 Astra, a new model with greater capacity to control your computer
That AI controls our PC completely is becoming less and less strange. We already saw what happens when we give control to an agent like OpenClaw. However, it is precisely this concept that OpenAI has delved into with its new AI model. One that, finally, leaves version 5.X behind and makes the leap to version 6. The laboratory has officially presented GPT-6 Astraits new model aimed at neither more nor less than automating desktop environments and streamlining advanced workflows. Do you remember the story of this laboratory’s massive purchase of Mac minis? It has a lot to do with all this.
The tool has already begun to be deployed among those who pay a ChatGPT subscription and through the company’s programming interface, in addition to being distributed on third-party platforms such as Microsoft Azure and AWS Bedrock. According to the first data, The performance jump compared to GPT-5.6 Sol is no anecdotebasically because Astra integrates native capabilities to interact directly with operating systems.
Let’s take a look at the numbers. In the synthetic tests of formal reasoning it has reached an overwhelming 99.9% in ARC-AGI-3, surpassing the efficiency of human action in 96% of the levels evaluated, according to data from the ARC Prize foundation.

As if that were not enough, the system achieves 98% success in FrontierMath Tier 4. So that we understand each other, these records reflect a greater mathematical capacity that seeks to extinguish once and for all the usual hallucinations when processing complex calculations or executing scientific simulation code.
Obviously, as we always say, all this data, despite being very striking, comes from the company itself. On the other hand, benchmarks are very specific tests and do not represent the capacity of the model in all situations. In any case, they do allow us to sense a substantial improvement over the predecessor GPTs.
When the algorithm takes control of your PC
In the area of interaction with graphical interfaces, things get interesting. The system records a score of 59.3% on the Agents’ Last Exam benchmark, beating Claude Opus 5’s 55.5% and, even better, achieving it with 65% fewer tokens generated.

In the simulations on the OSWorld 2.0 environment, the model reduces resolution time by 47% compared to the previous version. In practice, this means that AI can operate directly on complex applications such as office suites, browsers or printed circuit design tools such as KiCad.
For development environments, the Codex platform update increases execution speed by a factor of 1.9 within the Mind2Web test. Additionally, the company replaces the typical traditional compressed summary with an experimental system of structured notes that retains information between different context windows. This lifesaver preserves essential technical details during extended debugging sessions, preventing the model from losing track of changes applied in large software projects.
Dangerously bright
Let’s now talk about the most delicate part of the announcement. This is the area of cybersecurity, where the company has had to classify the system’s capabilities at the critical level after achieving 100% in ExploitBench in unrestricted tests. As if that were not enough, during recent vulnerability assessments, the algorithm single-handedly located and exploited two completely unknown zero-day security flaws.

Obviously, to avoid disasters, the commercial version will block the generation of malicious code, limiting its operations to defensive auditing and vulnerability mitigation tasks. It remains to be seen if we don’t see more attacks like that of Hugging Face.
OpenAI has designed specific tests to verify that autonomous agents do not exceed user-assigned limits. The good news is that, in isolated honeypot environments, the new model recorded 0% attempts to bypass access restrictions, compared to the worrying rate of 48.2% observed in its day with GPT-5.6 Sol. In addition, supervisory systems actively monitor the actions of the AI to abort any unauthorized operations before they interact with external networks.
However, the company’s technical reports openly admit that Astra’s written reasoning is more complex to audit than that of previous architectures due to the tremendous condensation of its intermediate steps. Although OpenAI rules out absolute opacity in its processes, it assumes that monitoring these logical chains represents a persistent challenge. Precisely for this reason, the company will keep automatic classifiers in production to interrupt processes if it detects signs of misalignment with respect to the initial instructions.
GPT-6 Astra pricing and availability
As we said at the beginning, the model will be progressively deployed among subscribers of ChatGPT’s Plus, Pro, Business and Enterprise plans, remaining deactivated by default in corporate environments so that administrators can expressly authorize its use. In the official API, the service already has a price. Specifically, we are talking about $10 per million tokens for entry and $50 per million for exit. Developers will also have a fast mode that doubles the response speed in exchange for, logically, doubling the computing cost.
