ไทย
← กลับไปยังไทม์ไลน์

โมเดลเดียวกัน สองคะแนน ARC-AGI-3 ห่างกัน 36 จุด

ยุค: ยุค LLM · ทศวรรษ 2020–ปัจจุบัน

วันที่ 2 กันยายน 2026 ARC Prize เผยแพร่ผล ARC-AGI-3 ที่ผ่านการตรวจสอบของ GPT-6 Astra ครอบคลุมการตั้งค่า harness สิบสองแบบ บนชุดประเมินกึ่งปิดชุดเดียวกัน โมเดลทำได้ 62.71% ด้วย harness มาตรฐาน ซึ่งยอมให้โมเดลเลือกเก็บบันทึกติดตัวไปเอง และทำได้ 98.55% ด้วย harness แบบ Provider Adapter ซึ่งรักษาสถานะการให้เหตุผลระหว่างคำขอและบีบอัดบทสนทนายาว Astra ใช้จำนวนการกระทำน้อยกว่าค่ามัธยฐานของมนุษย์ที่ถูกทดสอบใน 96% ของด่าน ตัวเลขนี้วัดโมเดลพร้อมกับนั่งร้านของมัน จึงบันทึกไว้ทั้งสองค่า

ARC Prize Foundation

แหล่งที่มาและหลักฐาน

  1. GPT-6 Astra — ARC-AGI Results

    ARC Prize Foundation · ประเภทแหล่งที่มา: institutional-archive

    ARC Prize records 62.71% with the Standard harness and 98.55% with the Provider Adapter harness for the same model on the same Semi-Private set, across twelve harness configurations.

    เผยแพร่: ·เข้าถึง:

  2. OpenAI’s GPT-6 Astra on ARC-AGI-3

    ARC Prize Foundation · ประเภทแหล่งที่มา: primary-announcement

    ARC Prize’s own analysis quotes 62.7% for $26K with the Standard harness and a best observed 99.9% for $19K with a Provider Adapter, and states that Astra used fewer actions than the median tested human on 96% of levels.

    เผยแพร่: ·เข้าถึง: