โมเดลเดียวกัน สองคะแนน ARC-AGI-3 ห่างกัน 36 จุด
ยุค: ยุค LLM · ทศวรรษ 2020–ปัจจุบัน
วันที่ 2 กันยายน 2026 ARC Prize เผยแพร่ผล ARC-AGI-3 ที่ผ่านการตรวจสอบของ GPT-6 Astra ครอบคลุมการตั้งค่า harness สิบสองแบบ บนชุดประเมินกึ่งปิดชุดเดียวกัน โมเดลทำได้ 62.71% ด้วย harness มาตรฐาน ซึ่งยอมให้โมเดลเลือกเก็บบันทึกติดตัวไปเอง และทำได้ 98.55% ด้วย harness แบบ Provider Adapter ซึ่งรักษาสถานะการให้เหตุผลระหว่างคำขอและบีบอัดบทสนทนายาว Astra ใช้จำนวนการกระทำน้อยกว่าค่ามัธยฐานของมนุษย์ที่ถูกทดสอบใน 96% ของด่าน ตัวเลขนี้วัดโมเดลพร้อมกับนั่งร้านของมัน จึงบันทึกไว้ทั้งสองค่า
ARC Prize Foundation
แหล่งที่มาและหลักฐาน
GPT-6 Astra — ARC-AGI Results
ARC Prize Foundation · ประเภทแหล่งที่มา: institutional-archive
ARC Prize records 62.71% with the Standard harness and 98.55% with the Provider Adapter harness for the same model on the same Semi-Private set, across twelve harness configurations.
เผยแพร่: ·เข้าถึง:
OpenAI’s GPT-6 Astra on ARC-AGI-3
ARC Prize Foundation · ประเภทแหล่งที่มา: primary-announcement
ARC Prize’s own analysis quotes 62.7% for $26K with the Standard harness and a best observed 99.9% for $19K with a Provider Adapter, and states that Astra used fewer actions than the median tested human on 96% of levels.
เผยแพร่: ·เข้าถึง: