Skip to content
Today

Signal

Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward

In shortEpoch AI reports that on the new FrontierMath Erdős, GPT-6 Astra was the only model to solve two of 68 open Erdős problems with Lean-verified proofs, on a budget of $300 per attempt.

Based on reporting from The Decoder; no matching primary source is listed in the Source Stack yet.

What happened

PRO-LONG is an agent framework developed by a third party that the ARC Prize team deployed early on as a red-teaming partner for ARC-AGI-3—that is, to systematically explore the limits of the benchmark. Epoch AI reports that on the new FrontierMath Erdős, GPT-6 Astra was the only model to solve two of 68 open Erdős problems with Lean-verified proofs, on a budget of $300 per attempt.

Why it matters

Evidence about GPT-6 Astra's reported benchmark performance can change model or system selection. Teams should reproduce the result on their own workload and compare failure cases before adopting it.

Who it affects

Researchers evaluating GPT-6 Astra's reported benchmark performance · Product teams testing GPT-6 Astra's reported benchmark performance

The bigger picture

GPT-6 Astra's reported benchmark performance fits a broader move from headline benchmarks toward reproducible evidence, disclosed failure modes, and tests that resemble real operating conditions.

What happens next

  • Watch for independent reproduction of GPT-6 Astra's reported benchmark performance, including failure cases and results on workloads that were not selected by the authors.

This fresh brief is based on concrete independent reporting; a matching official statement is not yet available.

Source Stack

Independent reporting

Put the news to work

Choose your next AI workflow

Reduce AI API costs without losing useful results

Measure workload, retries and accepted outputs before comparing models, batching or reusable prompts.

Check AI data handling before a team rollout

Work through account policies, retention, local records and access boundaries before sharing team data with an AI workflow.

Plan a local AI deployment you can verify

Check model routing, server access, external traffic and release changes before depending on a local AI workflow.

See what happened next · Compare verified API prices · Estimate a workload · Read the weekly index · Get future updates