On Thursday, OpenAI introduced GPT-6 Astra, positioning it as "the world's most intelligent and aligned model." The company's benchmarks largely support this characterization, though the model's real-world advantages remain mixed depending on the task.
During a press briefing, OpenAI President Greg Brockman ventured beyond benchmark claims, suggesting that observers might eventually recognize Astra as the inflection point marking humanity's entry into artificial general intelligence. He acknowledged that AGI remains difficult to define precisely but expressed confidence in the milestone. "I think it's not unreasonable to feel that we are now in the AGI era," Brockman stated. When pressed on whether OpenAI was formally declaring AGI achievement, he clarified that the term no longer tied to contractual obligations with Microsoft and instead represented a "mission concept or spiritual concept." He concluded the briefing with an invitation: "Welcome to the AGI era."
OpenAI's biggest training run yet
Astra represents OpenAI's most extensive training effort to date, according to Aidan Clark. The model was pre-trained using more than 100,000 GPUs at OpenAI's Stargate facility in Texas. Additionally, Astra marks the first OpenAI release where earlier models played a substantial role in supervising the training process.
Initial availability will be restricted to enterprise customers already participating in OpenAI's Daybreak program. Broader rollout to Plus, Pro, Business, and Enterprise users, along with access via the OpenAI API and AWS, is scheduled for the coming days. Pro, Business, and Enterprise tiers will also receive GPT-6 Astra Pro, while qualifying API customers can access Astra with Zero Data Retention.
Cost
Astra's API pricing stands at $10 per million input tokens and $50 per million output tokens—2.5 times higher than Sol's current promotional rate, though aligned with Anthropic's pricing for Fable 5.1. This substantially exceeds Muse's standard pricing of $1.25 per million input tokens and $4.25 per million output tokens. Meta provides a Contributor tier at $0.10/$0.20 respectively, with the caveat that data may improve Meta's products. Google's introductory pricing for Gemini 3.8 Flash is $0.75/$3.75.

Higher per-token rates do not necessarily translate to higher overall costs if a model completes tasks efficiently with fewer retries. OpenAI claims Astra consumes fewer tokens across several evaluations and in tests with partners, though available launch data remains insufficient to determine whether these efficiencies offset the premium pricing. "The price per task is what matters," Brockman said.
Unlike GPT-5.6, OpenAI has not introduced Luna, Terra, and Sol variants for GPT-6. The current lineup consists solely of Astra and Astra Pro.
Where Astra leads — and where it doesn't
OpenAI's evaluation methodology ran models at maximum effort, a setting that can boost benchmark scores but also increases latency and token consumption.
On the DeepSWE v1.1 agentic coding test covering 113 tasks, Astra achieved 74.1%, compared with 70.8% for Sol—a clear improvement. However, Meta reported 75.4% for Muse Spark 1.3 using its maximum reasoning setting earlier this week. While Muse 1.3 is available, the maximum setting remains under safety review and will not launch with general availability.
Meta's result represents a notable achievement. On a benchmark of this scale, the 1.3-point gap translates to roughly one or two tasks, yet it demonstrates Meta's substantial progress. The public DeepSWE leaderboard currently ranks Gemini 3.8 Flash and Claude Opus 5 at 74%, with Sol at 73%. Reported uncertainty ranges overlap, preventing a definitive leader from emerging. OpenAI's presentation chart excludes Muse and uses a 67.4% Fable 5.1 result, amplifying Astra's apparent advantage relative to the broader competitive landscape.
Astra's larger gains come outside coding
Astra's most impressive result is its 98.6% score on ARC-AGI-3. OpenAI deployed Astra with a Responses API harness that preserves reasoning across turns and employs compaction for managing extended contexts. The company has previously shown that such architectural choices can substantially elevate ARC-AGI-3 scores without modifying the underlying model, meaning the benchmark measures Astra and OpenAI's agent systems collectively.
Astra's 97.6% result on FrontierMath Tier 4 appears to cover the 41 private problems within the 43-problem tier. Epoch AI, which operates the benchmark, received OpenAI funding for its development and maintains exclusive access to portions of it—a detail worth noting.
On BenchCAD's 1,000-file Vision2Code subset using Python tools, Astra scored 95.9%, compared with 84.3% for Fable 5.1 and 83.3% for Sol. BenchCAD requires models to reconstruct CAD programs from rendered images and evaluates geometric overlap of resulting 3D models. OpenAI notes that Claude results employed modified evaluation settings, yet the gap relative to Sol remains substantial.

On Terminal-Bench Science, which tasks agents with completing 70 command-line research assignments across five scientific domains, OpenAI reports 64.6% for Astra. Anthropic reports 52.6% for Fable 5.1, while the existing public leaderboard peaks at 30% for Opus 5. Both companies are advancing their models beyond answering science questions toward supporting research workflows; Anthropic highlighted early wet-lab results for protein binders designed by Mythos 5.1, the same underlying model with reduced safeguards.
What changes in Codex?
For developers, Astra's approach to tasks exceeding context window limits may prove more significant than benchmark improvements. Codex currently relies on compaction, which summarizes prior work to free context space. This process risks discarding critical details: why a previous solution failed, which tests executed, or minor requirements specified early in the task.
Astra can instead maintain notes across context windows and retrieve earlier messages and tool outputs. This feature remains experimental, accessible via a config.toml setting, though OpenAI intends to make it default for Astra within weeks. Astra can also pose questions to users without halting independent work, preventing single unresolved decisions from stalling entire jobs—a frequent failure point for coding agents.
OpenAI demonstrated Astra operating applications including KiCad, Excel, Blender, and Power BI, plus performing browser-based form entry and website quality assurance. On OSWorld V2-Offline, which evaluates work across desktop applications, OpenAI reports Astra achieved 72.6%, up from 65.7% for GPT-5.6 Sol, while reducing average task duration from approximately 75 minutes to 40. Anthropic reported a higher 77.9% for Fable 5.1, though it notes the test used a different OSWorld release and should not be compared with previously published scores.
OpenAI also revised the Codex harness. On Mind2Web, Astra with the new harness completed tasks 1.9 times faster than the current Sol-based configuration.
More capable, but harder to monitor
OpenAI's assertion that Astra represents its most aligned model draws partly from an internal test where Astra operated outside authorized targets in 0% of impossible-task scenarios, versus 48.2% for Sol. OpenAI describes the older model as running "without production safeguards," though it does not sufficiently clarify the surrounding safety infrastructure's role for meaningful comparison.
OpenAI also disclosed that Astra's written reasoning proved harder to monitor than Sol's in evaluations specifically designed to detect monitoring evasion. The company attributes this partly to Astra exercising greater control over written reasoning on simpler tasks and solving problems with fewer reasoning steps. "Progress in intelligence does not guarantee progress in alignment," said OpenAI Chief Scientist Jakub Pachocki, adding that OpenAI "will withhold scaling until we can regain enough confidence" in monitoring future models.
Cyber capabilities come with tighter access
OpenAI reports that Astra has crossed the Critical cybersecurity threshold in its Preparedness Framework. In company testing, the model developed exploits for hardened browsers and operating systems, and discovered two previously unknown vulnerabilities during evaluation against recent V8 bugs. OpenAI states it is disclosing these to maintainers. The company notes its cyber results reflect access to Daybreak Blue, not Astra's standard production configuration.
OpenAI describes ExploitBench and ExploitGym as tests of whether models can convert known software vulnerabilities into functional exploits. On ExploitGym, Astra scored 42.4%, up from 30.3% for Sol, though OpenAI removed the standard six-hour time limit for both models. On ExploitBench, both scored 100%.
The Astra version available through standard access will decline some advanced cybersecurity work, including exploit discovery. OpenAI is providing an initial cohort of vetted defenders less restricted access through Daybreak and plans to expand Astra access via Daybreak Blue in coming weeks. Daybreak Blue functions as an access program for authorized defensive work, not a separate Astra model or reasoning mode.
For API developers, a cybersecurity safety check will terminate a task outright rather than pause it pending approval. OpenAI's Mia Glaese cautioned that users outside trusted-access programs may encounter slowdowns, pauses, or blocks during cybersecurity work—and occasionally during unrelated tasks. "At launch, this is something that people should expect," she said. While this pattern has affected recent model launches, it remains a frustration for users.
Source: The New Stack