نموذج واحد ودرجتان في ARC-AGI-3 تفصل بينهما ستّ وثلاثون نقطة
الحقبة: عصر النماذج اللغوية الكبيرة · عقد ٢٠٢٠–الحاضر
في ٢ سبتمبر ٢٠٢٦ نشرت ARC Prize نتائج ARC-AGI-3 المتحقَّق منها لنموذج GPT-6 Astra عبر اثنتي عشرة تهيئة لمنظومة التشغيل. وعلى مجموعة التقييم شبه الخاصة نفسها يسجّل النموذج ٦٢٫٧١٪ مع منظومة Standard التي تتيح له حمل ملاحظات يختار الاحتفاظ بها، و٩٨٫٥٥٪ مع منظومة Provider Adapter التي تحفظ حالة الاستدلال بين الطلبات وتضغط المحادثات الطويلة. واستعمل Astra إجراءات أقل من الإنسان الوسيط المختبَر في ٩٦٪ من المستويات. والرقم يقيس نموذجًا مع سقالته، ولهذا سُجّل الاثنان هنا.
ARC Prize Foundation
المصادر وما تسنده
GPT-6 Astra — ARC-AGI Results
ARC Prize Foundation · نوع المصدر: institutional-archive
ARC Prize records 62.71% with the Standard harness and 98.55% with the Provider Adapter harness for the same model on the same Semi-Private set, across twelve harness configurations.
تاريخ النشر: ·تاريخ الاطلاع:
OpenAI’s GPT-6 Astra on ARC-AGI-3
ARC Prize Foundation · نوع المصدر: primary-announcement
ARC Prize’s own analysis quotes 62.7% for $26K with the Standard harness and a best observed 99.9% for $19K with a Provider Adapter, and states that Astra used fewer actions than the median tested human on 96% of levels.
تاريخ النشر: ·تاريخ الاطلاع: