• AI Fire
  • Posts
  • 🤯 GPT-6 Astra Is Here. We Tested ALL the Little Things Everyone Misses

🤯 GPT-6 Astra Is Here. We Tested ALL the Little Things Everyone Misses

You’ll get our honest take on its performance, coding, reasoning, computer use, pricing, access, and the details you’ll want to know before using GPT-6 Astra.

TL;DR

GPT 6 Astra is a major upgrade focused on stronger reasoning, coding, computer use, and long multi-step workflows. The biggest change is that Astra can handle more of a task without needing constant instructions.

Its benchmark results are strong across ARC-AGI-3, FrontierMath, DeepSWE, Terminal-Bench, and AutomationBench. Real-world tests also show GPT 6 working inside codebases, browsers, Canva, Blender, Excel, and presentation workflows.

GPT 6 costs more than lighter models, so it makes the most sense for complex work where better task completion can reduce retries and manual fixes.

Key points

  • GPT 6 Astra scored 99.9% on ARC-AGI-3 in the benchmark shown.

  • GPT 6 is most useful for coding, research, computer use, and multi-step work.

  • Simple tasks usually don’t need Astra because cheaper models can handle them well.

Introduction

Everyone's screaming AGI for Astra.

Jensen Huang tweeted "AGI has arrived" the day it dropped. The benchmark charts are everywhere. 99.9% on this, 98% on that.

Cool. We ignored all of it. We ran Astra through the boring, real, everyday tasks. The ones you actually do at 11pm.

Some of what it did surprised us. One thing kind of annoyed us. Here's everything the hype missed. 👇

GPT 6 can handle almost anything you normally do on a computer, and it can do it very fast. How powerful is GPT 6, really?

🤖 Would you trust GPT 6 with real work?

Login or Subscribe to participate in polls.

I. So What Even Is GPT-6 Astra?

When OpenAI says Astra can "handle almost anything you normally do on a computer," that's not just marketing fluff, but it's also not the full picture.

GPT-6 Astra was released on September 3, 2026, initially as a limited preview for trusted partners, before rolling out to paid users the following day. The timing wasn't accidental, Anthropic had just shipped Claude Fable 5.1 on September 1, and OpenAI was clearly ready to answer.

→ OpenAI called Astra a model that "saturates ARC-AGI-3" and "sets a new frontier on computer and browser use, handling the most demanding professional work with unmatched speed, accuracy, and judgment."

what-is-gpt-6-astra

Bold words. So what does Astra actually do differently?

The real shift with GPT-6 is about agentic work, Astra doesn't wait for you to guide every step. Astra is "better at staying focused, adhering to task boundaries, understanding user intent, handling tedious tasks, and completing multi-step workflows."

OpenAI's president Greg Brockman said Astra could eventually be seen as the arrival of artificial general intelligence (AGI).

💡 Worth knowing: Following OpenAI's Hugging Face incident in July 2026, the company delayed the release to add more safeguards before shipping Astra. The rollout is staged, and for good reason, more autonomy means more risk.

II. GPT 6 Benchmarks: How Big is the Upgrade?

I usually don’t get too excited about benchmarks because every new model comes with a few strong numbers. But with GPT 6 Astra, some of the gaps are large enough that I wanted to look closer.

Key points

  • Reasoning: GPT 6 hits 99.9% on ARC-AGI-3 and about 98% on FrontierMath Tier 4.

  • Coding: Astra reaches about 74% on DeepSWE and 64% on Terminal-Bench Science.

  • Multi-step work: AutomationBench shows GPT 6 is especially strong on longer workflows.

1. Reasoning: The Big Numbers

ARC-AGI-3 is the number that stands out first. GPT 6 Astra scores 99.9%, while Claude Opus 5 reaches 30.2% and GPT-5.6 Sol gets only 7.8%.

arc-agi-3

This benchmark tests how well a model can find patterns and solve new problems. So a score of 99.9% suggests GPT 6 is handling this type of reasoning at a very different level from the previous generation.

FrontierMath Tier 4 shows a similar pattern. In the stronger settings shown in the chart, GPT 6 reaches around 98%, while GPT-5.6 Sol peaks at about 83%. Claude Fable 5 can reach around 90%, but at a much higher API cost.

frontiermath-tier-4-v2

I like this chart because it shows more than accuracy. GPT 6 is getting very high scores while keeping the benchmark cost much lower than several competing setups.

2. Coding: Closer Than OpenAI's Charts Suggest

On DeepSWE v1.1, GPT 6 reaches around 74% at its best setting. The gap is smaller here because GPT-5.6 Sol and Claude Fable 5 also get fairly close.

deepswe-v1-1

DeepSWE matters because it is closer to real software engineering than a short coding question. GPT 6 has to understand the task, change code, and keep track of context across a longer workflow.

Terminal-Bench 4.0 shows another strong result. GPT 6 reaches about 58% accuracy, while GPT-5.6 Sol gets around 37%. Claude Fable 5.1 also reaches about 56%, but at a higher API cost.

terminal-bench-4-0

The gap becomes clearer on Terminal-Bench Science 0.1. GPT 6 reaches about 64% resolution rate, while Claude Fable 5.1 tops out at around 53% in the chart.

terminal-bench-science-0-1

For me, these two Terminal-Bench results are more useful than a normal coding score. They test whether GPT 6 can work through a terminal, handle several steps, and keep going when the task doesn’t work perfectly on the first try.

3. Multi-Step Work: This Is Where Astra Really Shines

AutomationBench is the last chart I want to look at because it connects closely with how GPT 6 is meant to work. Astra reaches more than 40% accuracy at its strongest setting, while GPT-5.6 Sol stays below 20%.

automationbench

That doesn’t mean GPT 6 will complete every workflow perfectly on its own. Still, these benchmarks show that Astra is moving forward in the areas I care about most: harder reasoning, longer coding tasks, and workflows that need several steps.

I still wouldn’t choose GPT 6 just because it looks good on a leaderboard. The next question is more practical: how much does GPT 6 cost, and are the usage limits good enough for these longer workflows?

Benchmark Summary

Benchmark

GPT-6 Astra

Claude Fable 5.1

Winner

ARC-AGI-3

99.9%*

Not reported

Astra*

FrontierMath Tier 4

97.6%

87.8%

Astra ✅

Terminal-Bench Science

64.6%

52.6%

Astra ✅

DeepSWE

~74%

~63.6%

Astra (barely)

AutomationBench

41.4%

31.4%

Astra ✅

Humanity's Last Exam

57.2%

65.0%

Fable 5.1 ✅

Artificial Analysis Index

61

66

Fable 5.1 ✅

*The 99.9% ARC-AGI-3 score used OpenAI's stateful harness. Standard API calls score significantly lower (17–63% per ARC Prize's independent tests).

III. Pricing and Access: Who Can Actually Use It?

Strong benchmarks are nice, but pricing is where it gets real for builders. Here's what you need to know.

Key points

  • Pricing: GPT 6 Astra costs $10 input / $50 output per 1M tokens.

  • Usage limits: Access varies by plan, with higher tiers getting more room for longer workflows.

  • Access: GPT 6 is available through ChatGPT plans and the API, with rollout happening in stages.

1. API Price Tag

GPT-6 Astra is priced at $10 per million input tokens and $50 per million output tokens, with cached input at $1 and a Fast mode at 2x the rate. That's 2.5x GPT-5.6 Sol's current promotional pricing.

Model

Input

Output

GPT 6 Astra

$10 / 1M tokens

$50 / 1M tokens

GPT-5.6 Sol

$4 / 1M tokens

$20 / 1M tokens

Claude Fable 5.1

$10 / 1M tokens

$50 / 1M tokens

So on the surface, Astra and Fable 5.1 cost the same. However, Fable 5.1's cache reads cost only $0.25 per million tokens, versus Astra's $1.00.

2. Usage Limits by Plan

The standard version of GPT-6 Astra offers about half as many messages per five-hour window as GPT-5.6 Sol across all plans. The $200 Pro plan includes 200 messages per week for GPT-6 Pro, and the $100 Pro plan gets 50 per week.

OpenAI now publishes estimated Astra capacity per five-hour period: roughly 5–45 messages on Plus and Business Standard, 25–225 on Pro $100, and 100–900 on Pro $200.

Plan

Astra Access

Free

❌ Not available

Plus ($20/mo)

Limited Astra in Work and Codex

Pro $100/mo

Full allowance (50 GPT-6 Pro messages/week)

Pro $200/mo

Full allowance (200 GPT-6 Pro messages/week)

Business Standard

Limited Astra

Business Premium

Full allowance

Enterprise

Full allowance (admin must enable first)

API

$10/$50 per 1M tokens, pay as you go

If you're planning to run full workflows from start to end, that matters.

How useful was this GPT 6 breakdown?

Login or Subscribe to participate in polls.

IV. What Can GPT-6 Actually Do? (Real Builds)

This is honestly the most interesting part. People got access to Astra this week, and the things they built are harder to ignore than any benchmark chart.

1. Coding: An Entire FPS Map in 28 Minutes

Riley Brown gave Astra full computer access inside Codex and pointed it at a first-person shooter project called AFTERFLASH, asking it to list 10 meaningful improvements, pick the top 5, and do them. Astra picked 5 upgrades.

The run lasted 28 minutes and 16 seconds, changed 20 files, and passed 80 automated checks.

What I love about this one is that there's a session length. A file count. A list of decisions Astra made on its own. A test suite it ran. That's real evidence.

2. Browser-Based Game: From Prompt to Playable

In one test, a user gave GPT 6 a starting prompt and Astra built a 3D game that could run directly in the browser.

I find that more useful because you can start with the result you want, while GPT 6 handles more of the work in between.

Things get even more interesting when Astra has to leave the coding environment and use normal software.

3. Computer Use: Recreating a Portrait in Canva

One demo started with a simple request: “draw me inside Canva.” GPT 6 opened Canva, looked at the reference image, and worked directly inside the interface to recreate it.

Astra was actually looking at the screen and doing the actions itself. That’s why computer use matters. If GPT-6 can understand an interface and work inside it, you can give Astra many more types of jobs.

For harder tasks, though, Astra also needs to find information when the starting context isn’t enough.

4. 3D Modeling and Research

Sharif Shameem gave GPT 6 a more unusual job: recreate the Palace of Fine Arts in Blender. Astra searched for reference images and looked for extra information about the building before continuing the model.

I like this example GPT-6 found information because Astra needed that information to finish a bigger task.

That feels much closer to real work than asking for a summary and then doing everything else yourself.

5. Spreadsheets, Data, and Presentations

In OpenAI’s demo, GPT 6 worked directly with Excel and Power BI. Astra could read the data, make changes inside the file, and adjust the output based on the task.

You don’t have to ask GPT 6 how to create a chart and then go back to Excel to build it yourself. Astra can work inside the same environment where the data already lives.

That same workflow also fits documents and slides.

6. GPT-6 for Documents and Presentations

Simon Smith tested GPT 6 by giving Astra a CSV file about companies affected by AI. He asked GPT 6 to study the data and turn it into a complete presentation.

The first version was usable but a little plain, so Simon asked GPT 6 to improve the story, use more varied layouts, and add stronger visuals. He said the second version was about 95% ready to present, with only a few small edits left.

After looking at these examples, the main change feels clear to me. You can give GPT 6 a final goal, and Astra can handle much more of the work before giving the result back to you.

V. Is GPT-6 Astra Actually Worth It?

Use GPT-6 Astra when:

  • You're doing computer use: operating real apps, browsers, or design tools directly

  • You're working with advanced math or scientific research tasks

  • You need long, multi-step agent workflows (AutomationBench is where Astra really pulls ahead)

  • You're in Codex and want strong coding efficiency per dollar

Stick with a lighter model when:

  • You need a quick email, a rewrite, or a simple answer, Astra is overkill and costs more

  • You're running repeated-context agent pipelines where cache reads add up, Fable 5.1's $0.25/1M cache read rate makes it cheaper in that specific scenario

Conclusion

Astra is not a clean sweep. Its coding scores are a tie with Fable 5.1 on several benchmarks, the independent Intelligence Index shows no aggregate jump over its predecessor on general reasoning, and it costs about 2.5x more per token than GPT-5.6 Sol.

But for the tasks it was built for, such as computer use, automation, math, long agentic work, the results are genuinely different from what came before. That's the version of GPT-6 worth caring about.

If you are interested in other topics and how AI is transforming different aspects of our lives or even in making money using AI with more detailed, step-by-step guidance, you can find our other articles here:

Reply

or to participate.