同一モデルで 36 ポイント離れた二つの ARC-AGI-3 スコア
時代:LLM時代 · 2020年代~現在
2026年9月2日、ARC Prize は GPT-6 Astra の ARC-AGI-3 検証結果を十二通りの harness 構成にわたり公表した。同一の半非公開評価セットで、標準 harness では 62.71%、これはモデルが残すと決めた覚え書きを持ち越せる構成である。Provider Adapter harness では 98.55%、これは要求のあいだ推論状態を保持し、長い対話を圧縮する。Astra は 96% のレベルで、試験を受けた人間の中央値より少ない手数で解いた。この数値はモデルと足場を合わせて測ったものであり、だからこそ両方を記録する。
ARC Prize Foundation
出典と根拠
GPT-6 Astra — ARC-AGI Results
ARC Prize Foundation · 出典種別:institutional-archive
ARC Prize records 62.71% with the Standard harness and 98.55% with the Provider Adapter harness for the same model on the same Semi-Private set, across twelve harness configurations.
公開日: ·確認日:
OpenAI’s GPT-6 Astra on ARC-AGI-3
ARC Prize Foundation · 出典種別:primary-announcement
ARC Prize’s own analysis quotes 62.7% for $26K with the Standard harness and a best observed 99.9% for $19K with a Provider Adapter, and states that Astra used fewer actions than the median tested human on 96% of levels.
公開日: ·確認日: