GPT-6 Astra Guide to Features, Benchmarks, Price, and Access
GPT-6 Astra is OpenAI's new flagship model for computer use, professional work, coding, research, and advanced reasoning. Announced on September 3, 2026, it arrives with unusually strong headline results: 99.9% on ARC-AGI-3 with the Provider Adapter harness, 97.6% on FrontierMath Tier 4, 72.6% on OSWorld 2.0, and 100% on ExploitBench.
Those numbers explain the attention. The bigger story is what GPT-6 can do across a complete workflow. OpenAI presents Astra as a model that can browse, operate software, create polished documents, write and test code, analyze visual interfaces, and revise its approach while it works. It also crosses OpenAI's Critical threshold for cybersecurity capability, which brings stronger safeguards and tighter access controls.
This guide separates the launch headlines from the operational details. It covers the model's main features, official benchmark results, access, API pricing, safety limits, and what the release means for teams building AI-assisted workflows.
What Is GPT-6 Astra?
GPT-6 Astra is the first model in OpenAI's GPT-6 family. OpenAI describes it as its most capable and aligned model so far, with state-of-the-art performance across computer use, browsing, software engineering, cybersecurity, science, and professional work.
GPT-6 is presented as a model family. Astra is the initial release, focused on deep reasoning and reliable action across complex tasks.
Model Specs
The GPT-6 Astra API model supports a 1,050,000-token context window and up to 128,000 output tokens. Its knowledge cutoff is April 30, 2026. Developers can choose low, medium, high, xhigh, or max reasoning effort.
The model works with the Responses API and supports web search, file search, image generation, hosted shell, skills, computer use, tool search, and Model Context Protocol connections. That tool range makes it suitable for workflows where the model must gather information, act inside software, inspect results, and continue based on feedback.
Availability and Access
OpenAI is rolling out GPT-6 Astra gradually. The launch announcement says access begins with a limited group of organizations before expanding to ChatGPT Plus, Pro, Business, and Enterprise users. API access is also rolling out through the OpenAI API, Microsoft Azure, and Amazon Bedrock.
Availability can differ by account, workspace, region, and product surface during a staged release. Check the ChatGPT model picker or API dashboard for current account access.
Computer Use and Professional Work
The most important GPT-6 Astra feature is the combination of reasoning and action. Earlier models could describe a process or generate a draft. Astra is designed to carry more of the work through software interfaces, files, tools, and validation steps.

Official chart downloaded from OpenAI's GPT-6 Astra launch page.
Computer Use
GPT-6 Astra reaches 72.6% on OSWorld 2.0 while completing tasks in 47% less time than GPT-5.6 Sol, according to OpenAI. It also scores 92.7% on ScreenSpot-Pro, a benchmark that tests whether a model can identify interface targets accurately.
Computer use requires more than finding a button. An agent must plan several steps, notice interface changes, recover from errors, and verify the final state. OpenAI says Astra can respond to mid-turn steering, which makes long-running work easier to supervise.
The model's 41.4% score on AutomationBench shows that broad computer automation still has room to improve. That result is a useful counterweight to the strongest benchmark headlines. Teams should keep approvals, audit logs, and confirmation steps around consequential actions.
Documents and Design
OpenAI emphasizes professional output as a core capability. GPT-6 Astra can create and edit spreadsheets, presentations, documents, and visual assets. It is designed to evaluate layout and presentation quality along with the underlying content.
A model can research a market, organize findings in a spreadsheet, turn the analysis into a presentation, and revise the output after visual inspection. Human review remains important for calculations, brand requirements, and external claims.
Coding and Research
Astra combines a long context window with software engineering, browsing, shell access, and file tools. That combination supports larger codebase analysis, multi-file implementation, test execution, and research tasks that require evidence from several sources.
The same structure applies to business research. Reliable results still depend on source quality, current data, and explicit acceptance criteria. A strong model cannot repair missing or incorrect business data on its own.
Better Judgment
OpenAI also highlights judgment and clarification. GPT-6 Astra is designed to identify uncertainty, ask for missing information, and adapt when a user changes direction during execution. That behavior is essential for real workflows because many tasks begin with incomplete instructions.
The useful pattern is supervised autonomy: give the model a clear objective, scoped tools, and checkpoints for high-impact actions. This is especially relevant when AI agents automate Amazon operations.
GPT-6 Astra Benchmark Results
GPT-6 Astra posts leading results across reasoning, computer use, mathematics, and cybersecurity. The official numbers are impressive, but each score should be read with its evaluation conditions.

Official chart downloaded from OpenAI's GPT-6 Astra launch page.
| Benchmark | GPT-6 Astra | What It Measures |
|---|---|---|
| ARC-AGI-3, Provider Adapter | 99.9% | Novel abstract reasoning tasks |
| FrontierMath Tier 4 | 97.6% | Expert-level mathematical problems |
| ScreenSpot-Pro | 92.7% | Visual target recognition in interfaces |
| OSWorld 2.0 | 72.6% | Real computer-use tasks |
| AutomationBench | 41.4% | Broad automation performance |
| ExploitBench | 100% | Advanced cybersecurity tasks |
Abstract Reasoning
The 99.9% ARC-AGI-3 result was achieved through ARC Prize's Provider Adapter harness, which allows a customized inference setup. ARC Prize also reports a best score of 62.7% with its standard harness. Both results are valid, but they answer different questions.
The Provider Adapter score shows what Astra can achieve under an optimized setup. The standard harness provides a controlled comparison under common conditions.
On FrontierMath Tier 4, GPT-6 Astra scores 97.6%, up from 59.2% for GPT-5.6 Sol in OpenAI's comparison. These are difficult research-level math problems, so the improvement suggests a major increase in sustained technical reasoning.
Computer Use
ScreenSpot-Pro focuses on locating the correct visual element. OSWorld evaluates broader tasks in a real operating-system environment.
Astra's 92.7% ScreenSpot-Pro result indicates strong visual grounding. Its 72.6% OSWorld score shows progress in end-to-end execution and illustrates that accurate clicking is only one component of dependable computer use.
Coding and Science
OpenAI reports state-of-the-art performance in software engineering and science, alongside the 97.6% FrontierMath result. The release also introduces stronger long-horizon behavior through asynchronous tool calling and mid-turn steering.
For developers, the advantage is less fragmentation between planning, implementation, tool use, testing, and revision. Researchers can synthesize evidence across longer source collections. Outputs still need verification against original evidence.
Benchmark Caveats
Benchmarks measure performance under defined conditions. They do not guarantee the same result in a different interface, with different tools, or on private company data. Scores can also depend on inference settings, retry policies, scaffolding, and the amount of compute used at test time.
Evaluate GPT-6 Astra with a task set drawn from the intended workflow. Measure completion rate, correction rate, time saved, cost, and error severity. This reveals whether the model is useful for a particular organization.
Safety and Cybersecurity
GPT-6 Astra is the first OpenAI model to cross the company's Critical threshold for cybersecurity capability. That classification signals advanced ability in vulnerability research and exploitation. It also triggers stronger deployment restrictions and monitoring.
Critical Cyber Capability
OpenAI reports a 100% score on ExploitBench and 42.4% on ExploitGym. During benchmark work, the model identified two previously unknown vulnerabilities. These results point to real defensive value for finding and fixing software weaknesses.
The same capability can be misused. OpenAI says the launch model applies refusals to advanced exploit-development tasks, uses layered safeguards, and limits some access. The company also introduced Project Daybreak to support vetted defensive researchers.
Stronger Alignment
OpenAI's launch materials report improvements in refusal behavior, prompt-injection resistance, and tests for deceptive or misaligned action. The company evaluated the model across more than 54,000 internal Codex tasks as part of its safety work.

Official chart downloaded from OpenAI's GPT-6 Astra launch page. Lower is better; zero observed successes does not establish zero risk.
These results matter for tool-using systems because an agent can encounter hostile instructions inside websites, documents, or external content. Prompt-injection defenses reduce risk, but they do not eliminate it. Sensitive credentials, destructive tools, and external communications still require narrow permissions and human confirmation.
Monitorability Limits
OpenAI's safety overview also says GPT-6 Astra's chain-of-thought monitorability is lower than earlier models. In plain terms, internal reasoning traces may provide less reliable visibility into what the model is doing or why.
That limitation strengthens the case for outcome-based oversight. Teams should log tool calls, preserve source links, verify file changes, test code, and require explicit confirmation before publishing, purchasing, sending, or deleting. Operational evidence is more dependable than trying to infer safety from a hidden reasoning process.
GPT-6 Astra Pricing
GPT-6 Astra is a premium API model. OpenAI lists standard pricing at $10 per million input tokens and $50 per million output tokens. Cached input costs $1 per million tokens, while cache writes cost $12.50 per million tokens.
Standard API Pricing
The published API rates are:
| Token Type | Price per 1M Tokens |
|---|---|
| Input | $10.00 |
| Cached input | $1.00 |
| Cache write | $12.50 |
| Output | $50.00 |
Prompts longer than 272,000 input tokens receive higher rates for the full request: 2 times the input and cache pricing, and 1.5 times the output pricing. Batch and Flex processing are listed at half the standard rate where available.
The million-token context window is valuable, but filling it indiscriminately can be expensive. Retrieval, structured summaries, prompt caching, and smaller task-specific context packages can reduce cost and improve focus.
Fast Mode
OpenAI offers a Fast processing option that can provide up to twice the speed at twice the applicable price. It is useful when latency directly affects an interactive workflow or a time-sensitive production process.
Standard processing will usually be more economical for background research, report generation, batch analysis, and tasks that can run asynchronously. Teams should compare total completion time and correction cost rather than judging the model on token price alone.
The AGI Question
GPT-6 Astra's release immediately revived debate about artificial general intelligence. A near-perfect ARC-AGI-3 result, research-level math performance, strong computer use, and advanced cyber capability make that discussion understandable.
Progress, Not Proof
ARC Prize states that benchmark saturation does not prove AGI. The difference between Astra's 99.9% Provider Adapter result and 62.7% standard-harness result is another reason to avoid a simple label based on one score.
Astra works across more domains and longer task chains. It still depends on defined interfaces, tools, and data, and requires oversight for high-impact actions. The broader AGI question remains open.
What Changes Now
More work can now move from a prompt to a verified deliverable in one workflow. Research, planning, software operation, coding, analysis, and presentation can connect with fewer manual handoffs.
That shift will affect ecommerce as AI shopping changes product discovery. Teams need machine-readable product data and clear action policies. OpenClaw for Amazon sellers illustrates how tool access and commerce context can focus a general model on specific tasks.
Conclusion
GPT-6 Astra is a significant release because its improvements extend beyond chat. It combines strong reasoning with computer use, professional output, long context, software tools, and flexible steering. Its benchmark results set new marks in several categories, while the standard-harness ARC score, AutomationBench result, cyber classification, and monitorability limits provide essential context.
For organizations, the next step is practical evaluation. Test Astra on real tasks, constrain its permissions, measure end-to-end outcomes, and supply trusted business data. The model can provide the reasoning layer, but the quality of a production workflow also depends on the context connected to it.
Add Ecommerce Context to GPT-6
Powerful reasoning becomes more useful when it can access structured, current ecommerce context. Nexscope provides supported product, keyword, review, competitor, pricing, sales, and trend data that can connect to a team's own ChatGPT or OpenAI workflow through MCP or REST API. ChatGPT connector availability depends on the plan and workspace settings.
Connect Ecommerce Data to GPT-6
Bring supported product, keyword, review, competitor, pricing, sales, and trend data into your own ChatGPT or OpenAI workflow through MCP or REST API.
Discover Nexscope Ecommerce APIs →Frequently Asked Questions
What is GPT-6 Astra?
GPT-6 Astra is the first model in OpenAI's GPT-6 family. It is designed for reasoning, computer use, browsing, coding, science, cybersecurity, and professional work. The API supports a 1,050,000-token context window, up to 128,000 output tokens, multiple reasoning levels, and tools including search, shell, computer use, and MCP.
When was GPT-6 Astra released?
OpenAI announced GPT-6 Astra on September 3, 2026. The staged rollout expands from selected organizations to ChatGPT Plus, Pro, Business, and Enterprise users, with API access through OpenAI, Microsoft Azure, and Amazon Bedrock. Timing can vary by account and workspace.
How much does GPT-6 Astra cost?
Standard API pricing is $10 per million input tokens and $50 per million output tokens. Cached input costs $1 per million, and cache writes cost $12.50 per million. Requests above 272,000 input tokens use higher long-context rates. Fast processing offers up to twice the speed at twice the applicable price.
Is GPT-6 Astra available in ChatGPT?
OpenAI says GPT-6 Astra is rolling out to ChatGPT Plus, Pro, Business, and Enterprise. It may appear at different times across accounts, workspaces, or regions. Check the model picker and workspace controls for current access.
What is the GPT-6 Astra context window?
The API documentation lists a 1,050,000-token context window and a maximum output of 128,000 tokens. This can support large codebases and document collections. Prompts above 272,000 input tokens cost more, so retrieval and caching remain important.
Does GPT-6 Astra support computer use?
Yes. OpenAI reports 72.6% on OSWorld 2.0 and 92.7% on ScreenSpot-Pro. Astra can interpret interfaces and continue through multi-step software tasks. Consequential actions should still use scoped permissions, logs, checkpoints, and human confirmation.
Does GPT-6 Astra support MCP?
Yes. GPT-6 Astra supports remote MCP through the OpenAI Responses API. External tools and data sources can supply approved context or actions. Inside ChatGPT, custom MCP capabilities depend on the plan, developer-mode access, workspace configuration, and administrator settings.
Is GPT-6 Astra AGI?
There is no settled answer. Astra scored 99.9% on ARC-AGI-3 with the Provider Adapter harness, while ARC Prize reports 62.7% with the standard harness and says benchmark saturation does not prove AGI. The model still depends on tools, data, evaluation conditions, permissions, and oversight.
Sources
- OpenAI. (2026). GPT-6 Astra: A New Generation of Intelligence. Retrieved from openai.com
- OpenAI. (2026). GPT-6 Astra Model Documentation. Retrieved from developers.openai.com
- OpenAI. (2026). Safety Overview: GPT-6 Astra. Retrieved from openai.com
- OpenAI. (2026). GPT-6 Astra System Card. Retrieved from deploymentsafety.openai.com
- ARC Prize Foundation. (2026). OpenAI's GPT-6 Astra on ARC-AGI-3. Retrieved from arcprize.org
