ARC Prize ran GPT-6 Astra twice. Its own harness returned 62.7%, OpenAI’s returned 99.9%, and the benchmark’s authors say they are not claiming AGI. OpenAI then revised five published metrics after launch. OpenAI declared the AGI era on the strength of a 99.9% score. Run the same model through the …


