同一个模型的两个 ARC-AGI-3 分数,相差 36 个百分点
时代:大模型时代 · 2020年代—现在
2026年9月2日,ARC Prize 公布了对 GPT-6 Astra 的 ARC-AGI-3 核验结果,覆盖十二种 harness 配置。在同一个半私有评测集上,标准 harness 下模型得 62.71%,该 harness 允许模型自行选择带走哪些笔记;Provider Adapter harness 下得 98.55%,它在请求之间保留推理状态并对长对话做压缩。Astra 在 96% 的关卡上使用的动作数少于受测人类的中位数。这个分数衡量的是模型连同它的脚手架,因此两个数字都记录在此。
ARC Prize Foundation
来源与支持内容
GPT-6 Astra — ARC-AGI Results
ARC Prize Foundation · 来源类型:institutional-archive
ARC Prize records 62.71% with the Standard harness and 98.55% with the Provider Adapter harness for the same model on the same Semi-Private set, across twelve harness configurations.
发布日期: ·访问日期:
OpenAI’s GPT-6 Astra on ARC-AGI-3
ARC Prize Foundation · 来源类型:primary-announcement
ARC Prize’s own analysis quotes 62.7% for $26K with the Standard harness and a best observed 99.9% for $19K with a Provider Adapter, and states that Astra used fewer actions than the median tested human on 96% of levels.
发布日期: ·访问日期: