Skip to content

OpenAI's GPT-6 Astra Breaks ARC-AGI-3 Records, Surpasses Human Efficiency and Develops Its Own Symbolic Language

Sep 04, 2026
ARC Prize
Article image for OpenAI's GPT-6 Astra Breaks ARC-AGI-3 Records, Surpasses Human Efficiency and Develops Its Own Symbolic Language

Summary

OpenAI's GPT-6 Astra shatters AI benchmarks by scoring 99.9% on ARC-AGI-3, surpassing human efficiency on 96% of levels while spontaneously developing its own symbolic language to navigate complex environments.

Key Points

  • OpenAI's GPT-6 Astra achieves state-of-the-art scores on the ARC-AGI-3 benchmark, reaching 62.7% with a Standard harness for $26K and an impressive 99.9% with a Provider Adapter harness for $19K on the Semi-Private evaluation.
  • Astra surpasses human action efficiency on ARC-AGI-3, using fewer actions than the median human on 96% of levels and averaging 51.7% fewer actions per level, marking a significant milestone in agentic AI performance.
  • A standout behavior observed in Astra is its ability to convert unfamiliar game environments into compact symbolic world models, spontaneously developing custom algebraic notation and domain-specific shorthand to track game state and plan multi-step actions.

Tags

Read Original Article