"GPT-6 Astra surpasses the human baseline in action efficiency on ARC-AGI-3. It used fewer actions than the median tested human on 96% of levels."
They go on to say "GPT-6 Astra scores 62.7% for $26K on ARC-AGI-3 Semi-Private with our Standard harness, and 99.9% for $19K with a Provider Adapter harness", and "A key behavior observed in GPT-6 Astra was its ability to turn unfamiliar environments into compact symbolic world models. It represented game mechanics as logical rules and developed its own domain-specific language shorthand to track state and plan actions."
If you're unfamiliar with the ARC-AGI tests, they're pattern-matching tests that are intended to be easy for humans but hard for AI despite AI being able to solve math problems, etc, that are hard for humans. It's intended to give a more realistic idea how close AI models are to surpassing human intelligence at more or less everything, not just specific tasks that are hard for humans. This threshold is called "artificial general intelligence" (AGI) which is not exactly an intuitive term (but math and science and the AI field are full of terms that are not intuitive -- can you come up with a better one?). ARC-AGI-3 is the 3rd version of the test, because AI keeps getting better and they keep having to make the test harder.
"These environments only contain core knowledge priors and are difficulty-calibrated through controlled testing with human participants. Humans can solve 100% of the environments."
"The goal of the ARC-AGI series is to measure the 'residual gap' between current artificial intelligence and AGI. We define AGI as a system's ability to acquire any skill a human can, as efficiently as a human can."
If you're wondering what the explanation is for the dollar amounts in the description above, they say: "For a cost comparison, during our controlled testing, human participants were paid $115 per 90-minute session, plus $5 per game completed. Participants attempted approximately nine games per session, roughly $12.78 per attempted game before bonuses."
"Most of this fee pays for the participant's time and willingness to take the test, rather than the energy their brain uses (a closer proxy to compare with AI). If we look at only the brain's energy, and price it as electricity, the estimate drops to about 0.6 cents per session, or 0.067 cents per game attempted."
"Beyond the scores, Astra's replays show how it turns unfamiliar game mechanics into useful working models. Three findings stood out: the compact algebraic notation it develops, its action efficiency compared with humans, and the custom tools it builds."
If you're wondering what the bit about "action efficiency" is all about, they say: "For each level, we defined the 'human baseline' using the median action count among players who completed it. This gives us a reference for comparing human and AI performance. An AI that needs more actions is less action-efficient, while one that needs fewer actions is more action-efficient."
"Before we launched ARC-AGI-3, we hypothesized that action efficiency would remain a dividing line between humans and AI. We anticipated that even when an AI solved an environment, it might require substantially more exploration (actions) than a person. That remains true of brute-force approaches, but frontier AI shows a more binary-like pattern. Once frontier AI 'understands' the mechanics, it generally executes within the range of human efficiency."
"Astra's results show that it needed fewer interactions than the human baseline to execute a solution."
There's also some stuff about "harnesses". They say: "The Standard harness enables a model to carry forward notes it chooses to keep with it throughout the environment", and "The Provider Adapter harness preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work."
OpenAI's GPT-6 Astra on ARC-AGI-3
#solidstatelife #ai #genai #llms #agi
OpenAI's GPT-6 Astra on ARC-AGI-3 | ARC Prize
OpenAI's GPT-6 Astra reaches state-of-the-art results on ARC-AGI-3ARC Prize