“Frontier models score under 1% on ARC-AGI-3” was true on 25 March 2026. By 24 July the best published number was 30.16%. By 3 September it was 62.7% under the shared harness and 99.9% under the vendor’s own. Most of the pages that quote the first figure have not been updated since it was the only one.
Here are the ARC-AGI-3 scores in date order, from the primary sources, with the set and the harness attached to each, because a score without those two labels is not comparable to anything.
What the ARC-AGI-3 scores were at launch in March 2026
The launch post gave two numbers: humans at 100% of environments and frontier AI at 0.51%. The technical report broke the second one out, on the semi private set, as the leaderboard at release:
| System | Score |
|---|---|
| Anthropic Opus 4.6 (Max) | 0.50% |
| Google Gemini 3.1 Pro Preview | 0.40% |
| OpenAI GPT 5.4 (High) | 0.20% |
| xAI Grok-4.20 | 0.10% |
Those are shared harness, semi private set numbers, the only kind the official leaderboard reports. They were low for a reason that the scoring explains better than the models do. A level counts only for the efficiency with which it was completed, squared against a first time human, so a system that finished levels by taking ten times the human’s moves earned a hundredth of the credit for each, and a system that did not finish them earned nothing. Under that rule, half a per cent was the sum of a great many nearly zero levels. The benchmark was built to be scored on efficiency rather than completion, which is the main thing separating it from ARC-AGI-2, and at launch no frontier system was efficient enough on enough levels to register above a rounding error.
How the ARC-AGI-3 scores moved between April and August
The results page lists the semi private, shared harness results that followed, each with its date:
| Date | System | Score |
|---|---|---|
| 9 July 2026 | GPT-5.6 | 7.78% |
| 16 July 2026 | Grok 4.5 | 0.32% |
| 24 July 2026 | Claude Opus 5 | 30.16% |
| 11 August 2026 | Grok 4.6 | 2.11% |
Four months after launch, one system had gone from half a point to thirty. That is the first place the “under 1%” line stopped being true, and it happened in July.
The same weeks produced a different kind of number. On 21 August NVIDIA reported a 100.00 RHAE on the public set with an agent architecture wrapped around Claude Opus 5, and said in the same post that “these results cover the 25-environment ARC-AGI-3 public set” and “are not results on the semi-private or fully private competition sets.” The public set is the 25 games anyone can download as source. A perfect score there is a development milestone, and NVIDIA labelled it as one. The post gives the detail that makes it checkable: all 183 levels solved, in 6,624 environment actions, with Claude Opus 5 as the model inside the architecture.
The competition track was moving on its own calendar. Its first milestone closed on 30 June and went to three entries running open weight models locally, a 27B and two 31Bs, all required by the rules to be open sourced under CC0 or MIT-0. The second milestone closes on 30 September, final submissions on 2 November, results on 4 December. None of those entries appear on the leaderboard above, because the leaderboard measures systems served behind a general purpose API and the competition measures whatever a team can build and give away.
What the ARC-AGI-3 scores are as of September 2026
On 3 September ARC Prize published GPT-6 Astra, and for the first time reported two conditions for one model on the semi private set:
| Harness | Score | Cost |
|---|---|---|
| Standard, shared with every other system | 62.7% | $26,098 |
| OpenAI’s Provider Adapter | 99.9% | $18,817 |
The Provider Adapter, in ARC Prize’s words, “preserves opaque reasoning state between requests and uses compaction for longer conversations.” The 37 point gap between the rows is the difference between a model that starts each turn cold and the same model allowed to keep its working memory, and it is the number this benchmark is currently teaching everyone.
Two footnotes belong on those rows. ARC Prize’s results page lists the Provider Adapter figure as 99.95% where the post rounds it to 99.9%, so both spellings are in circulation for the same run. And the post commits to reporting both harness conditions separately from now on, each one labelled, which means the next set of numbers will arrive as pairs and should be quoted as pairs.

Top published ARC-AGI-3 score on the semi-private set, by date. Yellow: the shared harness number that compares across systems. Black: the same model under its vendor’s own harness.
So the honest one line summary in September is: 62.7% shared harness, 99.9% vendor harness, both on the semi private set, and 100% on the public development set from more than one team.
Which ARC-AGI-3 score to quote, and when it goes stale
Quote the shared harness, semi private number when comparing systems, because it is the only one where every system ran the same code on the same games. Quote the vendor harness number when the question is what a model can do with its own scaffolding, and say so. Never quote a public set number as a benchmark result, and never quote any of them without the date.
Then expect the numbers to move. Between March and September the shared harness top score went 0.51, 7.78, 30.16, 62.7. Nothing about that curve suggests it has finished, and ARC Prize’s own results page is updated as systems are evaluated. This piece is on the quarterly re-check list for that reason, and without it, it will be wrong within a quarter in the same way the launch coverage is wrong now.
Two things will not move. The $700,000 grand prize is unclaimed until an agent scores 100% on the competition’s fully private set under the competition’s rules, and the first milestone of that competition went to entries running 27B and 31B open weight models locally. The leaderboard measures frontier APIs. The prize measures whatever anyone can build and open source, and those are the two curves worth watching separately.
