Italiano
← Torna alla cronologia

Un solo modello, due punteggi ARC-AGI-3 distanti 36 punti

Era: L’Era degli LLM · anni 2020–Presente

Il 2 settembre 2026 ARC Prize ha pubblicato risultati ARC-AGI-3 verificati per GPT-6 Astra su dodici configurazioni di harness. Sullo stesso insieme di valutazione semi-privato il modello ottiene 62,71% con il harness standard, che gli consente di portarsi avanti gli appunti che sceglie di tenere, e 98,55% con un harness Provider Adapter, che conserva lo stato di ragionamento fra le richieste e comprime le conversazioni lunghe. Astra ha usato meno azioni della mediana degli umani testati nel 96% dei livelli. La cifra misura un modello insieme alla sua impalcatura, ed è per questo che qui compaiono entrambe.

ARC Prize Foundation

Fonti e riscontri

  1. GPT-6 Astra — ARC-AGI Results

    ARC Prize Foundation · Tipo di fonte: institutional-archive

    ARC Prize records 62.71% with the Standard harness and 98.55% with the Provider Adapter harness for the same model on the same Semi-Private set, across twelve harness configurations.

    Pubblicato: ·Consultato:

  2. OpenAI’s GPT-6 Astra on ARC-AGI-3

    ARC Prize Foundation · Tipo di fonte: primary-announcement

    ARC Prize’s own analysis quotes 62.7% for $26K with the Standard harness and a best observed 99.9% for $19K with a Provider Adapter, and states that Astra used fewer actions than the median tested human on 96% of levels.

    Pubblicato: ·Consultato: