What is ARC-AGI-3?
ARC-AGI-3 is a benchmark from the ARC Prize Foundation that tests whether an AI agent can learn a task it has never seen: it is placed in a small turn-based game with no instructions and no stated goal, and scored on how efficiently it works the game out compared with a person. At launch in March 2026 humans scored 100 % and frontier models 0.51 %.
What the test actually asks
The first two ARC benchmarks were static puzzles: a few examples of a grid transformation, then a new grid to complete. Models learned to do those. ARC-AGI-3, launched on 25 March 2026, changes the shape of the question. The agent is dropped into one of hundreds of hand-made, turn-based environments and told nothing - not the rules, not what counts as winning, not which of its actions matter. It has to explore, notice what changes when it acts, form a theory, plan several moves ahead and correct itself when the theory is wrong.
The score is not "did it finish" but "how efficiently did it learn", measured against people who were given the same games cold. That is why the foundation writes a human score of 100 % and, at launch, a frontier score of 0.51 %: models that pass bar exams and write working code still could not work out a game a child solves in a few minutes. The foundation's own line is that as long as that gap exists, there is no AGI, and it has put two million dollars of prize money behind closing it.
The settings lesson: how OpenAI tripled its score without training
In August 2026 OpenAI reported that it had tripled its ARC-AGI-3 score by changing two inference settings - how the model was configured to think through a problem - with no retraining and no new model. The base was low and the post did not claim the effect transfers to other tasks, so this is not a story about ARC-AGI-3 being solved.
It is a story about defaults. The capability was already inside a model people were paying for, switched off by configuration nobody had revisited. The same thing happens in ordinary deployments: a support triage bot or a drafting assistant is judged "not good enough" on settings tuned for cost and speed, and nobody tests the alternatives on real work before deciding to rebuild or to give up.
What a benchmark score tells a buyer, and what it does not
ARC-AGI-3 measures one narrow thing well: learning an abstract task from scratch. It says nothing about whether a model reads your contracts accurately, follows a fourteen-step returns procedure, or stays polite on the four-hundredth ticket of the day. Those are different skills, and models that are close on ARC can be far apart on them.
So read benchmark news the way you would read a lab result for a car: interesting, occasionally decisive, never a substitute for a test drive. The only score that matters for a purchase is the one on your own tasks, with your own data, measured before and after a change. That is the test we run in a diagnostic, and it is the one most vendors would rather you did not.
Related questions
- How is ARC-AGI-3 different from ARC-AGI-1 and ARC-AGI-2?
- The first two are static: a handful of examples, then a puzzle to complete in one answer. ARC-AGI-3 is interactive - the agent acts inside a game over many turns, discovers the rules itself and is scored on how efficiently it learns, not only on whether it finishes.
- Who runs ARC-AGI-3?
- The ARC Prize Foundation, a non-profit co-founded by François Chollet, who created the original ARC in 2019, and Mike Knoop, a co-founder of Zapier. It publishes the benchmark, the human baselines and the prize competition.
- Does a higher ARC-AGI-3 score make a model better for my company?
- Not on its own. The benchmark measures abstract learning in unfamiliar games; a business task is mostly familiar work done reliably at volume. Test candidate models on a sample of your real cases and compare the results - that number transfers, a benchmark score does not.
- Is passing ARC-AGI-3 the same as AGI?
- The foundation is careful to say the opposite: a 100 % score means an agent matches human learning efficiency on these games, which it treats as a necessary condition, not proof of general intelligence.
Related
Free AI Diagnostic
The free diagnostic runs the test that transfers: your own cases, your own data, a before-and-after number for each candidate model.
Start the free diagnosticStarts immediately in the browser.
- Fee
- Free
- Length
- 15 minutes
You keep the ranked list of candidates either way.