OpenAI's GPT-6 Astra Breaks ARC-AGI-3 Records, Surpasses Human Efficiency and Develops Its Own Symbolic Language
Summary
OpenAI's GPT-6 Astra shatters AI benchmarks by scoring 99.9% on ARC-AGI-3, surpassing human efficiency on 96% of levels while spontaneously developing its own symbolic language to navigate complex environments.
Key Points
- OpenAI's GPT-6 Astra achieves state-of-the-art scores on the ARC-AGI-3 benchmark, reaching 62.7% with a Standard harness for $26K and an impressive 99.9% with a Provider Adapter harness for $19K on the Semi-Private evaluation.
- Astra surpasses human action efficiency on ARC-AGI-3, using fewer actions than the median human on 96% of levels and averaging 51.7% fewer actions per level, marking a significant milestone in agentic AI performance.
- A standout behavior observed in Astra is its ability to convert unfamiliar game environments into compact symbolic world models, spontaneously developing custom algebraic notation and domain-specific shorthand to track game state and plan multi-step actions.