同一個模型的兩個 ARC-AGI-3 分數,相差 36 個百分點
時代:大模型時代 · 2020年代—現在
2026年9月2日,ARC Prize 公布了對 GPT-6 Astra 的 ARC-AGI-3 核驗結果,涵蓋十二種 harness 配置。在同一個半私有評測集上,標準 harness 下模型得 62.71%,該 harness 允許模型自行選擇帶走哪些筆記;Provider Adapter harness 下得 98.55%,它在請求之間保留推理狀態並對長對話做壓縮。Astra 在 96% 的關卡上使用的動作數少於受測人類的中位數。這個分數衡量的是模型連同它的鷹架,因此兩個數字都記錄在此。
ARC Prize Foundation
來源與支持內容
GPT-6 Astra — ARC-AGI Results
ARC Prize Foundation · 來源類型:institutional-archive
ARC Prize records 62.71% with the Standard harness and 98.55% with the Provider Adapter harness for the same model on the same Semi-Private set, across twelve harness configurations.
發布日期: ·存取日期:
OpenAI’s GPT-6 Astra on ARC-AGI-3
ARC Prize Foundation · 來源類型:primary-announcement
ARC Prize’s own analysis quotes 62.7% for $26K with the Standard harness and a best observed 99.9% for $19K with a Provider Adapter, and states that Astra used fewer actions than the median tested human on 96% of levels.
發布日期: ·存取日期: