같은 모델, 36포인트 차이 나는 두 ARC-AGI-3 점수
시대: 대규모 언어 모델 시대 · 2020년대~현재
2026년 9월 2일 ARC Prize는 GPT-6 Astra의 ARC-AGI-3 검증 결과를 열두 가지 harness 구성에 걸쳐 공개했다. 같은 반비공개 평가 집합에서 모델은 표준 harness로 62.71%를 얻는다. 이 harness는 모델이 남기기로 고른 메모를 이어갈 수 있게 한다. Provider Adapter harness로는 98.55%를 얻으며, 이는 요청 사이에 추론 상태를 보존하고 긴 대화를 압축한다. Astra는 96%의 레벨에서 시험받은 사람의 중앙값보다 적은 동작을 썼다. 이 수치는 모델과 그 발판을 함께 잰 것이므로 둘 다 기록한다.
ARC Prize Foundation
출처와 근거
GPT-6 Astra — ARC-AGI Results
ARC Prize Foundation · 출처 유형: institutional-archive
ARC Prize records 62.71% with the Standard harness and 98.55% with the Provider Adapter harness for the same model on the same Semi-Private set, across twelve harness configurations.
게시일: ·접근일:
OpenAI’s GPT-6 Astra on ARC-AGI-3
ARC Prize Foundation · 출처 유형: primary-announcement
ARC Prize’s own analysis quotes 62.7% for $26K with the Standard harness and a best observed 99.9% for $19K with a Provider Adapter, and states that Astra used fewer actions than the median tested human on 96% of levels.
게시일: ·접근일: