Monday, September 7, 2026HotTea verified storyVerified 2:39 AM PDT
← Back to the Monday, September 7, 2026 edition

Astra scored 99.9 percent with OpenAI's hidden state and 62.7 percent without it

HotTea covered GPT-6 Astra's release. Since then, outside evaluators have tested the model under different conditions. ARC Prize measured 62.7 percent on ARC-AGI-3 with its standard test setup. That setup lets the model carry visible notes between steps. A provider adapter preserved OpenAI's opaque reasoning state and used compaction. That run reached 99.9 percent. ARC Prize called both results state of the art. Artificial Analysis found a different split. Astra scored 61 on its general Intelligence Index, the same as GPT-5.6 Sol. Its Coding Agent Index rose to 67, up from 65 for Sol.

Verified 2:39 AM PDT · 3 original sources

The 37.2-point ARC gap does not point to one cause by itself. The runs used different context systems, and their best scores came at different reasoning levels. The benchmark also uses bounded game environments, not open-ended work. Artificial Analysis says Astra's token price is 2.5 times Sol's. Astra used fewer tokens, but still cost 75 percent more per task on the general index.

Teams should compare Astra and its predecessor on complete jobs with the same tools, context rules and stopping conditions. Watch cost per successful task, not only cost per token. More provider-neutral tests will show which gains come from the model and which depend on OpenAI's context system.

Audit the story

Original sources

Company claims remain company claims. Follow the reporting and judge the evidence directly.

  1. OpenAIGPT-6 Astra: A new generation of intelligence ↗
  2. ARC PrizeOpenAI's GPT-6 Astra on ARC-AGI-3 ↗
  3. Artificial AnalysisBenchmarking GPT-6 Astra ↗

Continue the edition

Read the full edition.

Read the full editionListen to the daily audio →