Astra scored 99.9 percent with OpenAI's hidden state and 62.7 percent without it
HotTea covered GPT-6 Astra's release. Since then, outside evaluators have tested the model under different conditions. ARC Prize measured 62.7 percent on ARC-AGI-3 with its standard test setup. That setup lets the model carry visible notes between steps. A provider adapter preserved OpenAI's opaque reasoning state and used compaction. That run reached 99.9 percent. ARC Prize called both results state of the art. Artificial Analysis found a different split. Astra scored 61 on its general Intelligence Index, the same as GPT-5.6 Sol. Its Coding Agent Index rose to 67, up from 65 for Sol.
The 37.2-point ARC gap does not point to one cause by itself. The runs used different context systems, and their best scores came at different reasoning levels. The benchmark also uses bounded game environments, not open-ended work. Artificial Analysis says Astra's token price is 2.5 times Sol's. Astra used fewer tokens, but still cost 75 percent more per task on the general index.
Read article →