English
← Back to timeline

One Model, Two ARC-AGI-3 Scores Thirty-Six Points Apart

Era: The LLM Era · 2020s–Present

On September 2, 2026 ARC Prize published verified ARC-AGI-3 results for GPT-6 Astra across twelve harness configurations. On the same Semi-Private evaluation set the model scores 62.71% with the Standard harness, which lets it carry forward notes it chooses to keep, and 98.55% with a Provider Adapter harness, which preserves reasoning state between requests and compacts long conversations. Astra used fewer actions than the median tested human on 96% of levels. The figure measures a model together with its scaffolding, which is why both are recorded here.

ARC Prize Foundation

Sources and support

  1. GPT-6 Astra — ARC-AGI Results

    ARC Prize Foundation · Source type: institutional-archive

    ARC Prize records 62.71% with the Standard harness and 98.55% with the Provider Adapter harness for the same model on the same Semi-Private set, across twelve harness configurations.

    Published: ·Accessed:

  2. OpenAI’s GPT-6 Astra on ARC-AGI-3

    ARC Prize Foundation · Source type: primary-announcement

    ARC Prize’s own analysis quotes 62.7% for $26K with the Standard harness and a best observed 99.9% for $19K with a Provider Adapter, and states that Astra used fewer actions than the median tested human on 96% of levels.

    Published: ·Accessed: