- GPT-5.6 Sol ARC-AGI-3 Benchmark Settings: What High vs Max Show
- The score comparison that matters: Public versus Semi-Private
- What the results show across Sol, Terra, and Luna
- What the ARC-AGI-3 benchmark analysis actually measures
- What earlier replay research can and can't explain about Sol's gain
- Checking a benchmark claim before using it
GPT-5.6 Sol ARC-AGI-3 Benchmark Settings: What High vs Max Show
Three weeks ago, ARC Prize published a results page comparing GPT-5.6 Sol ARC-AGI-3 benchmark settings across five reasoning-effort levels, and one line in that table is worth slowing down on. Within the same reported Sol variant, moving the effort setting from High to Max lifted the ARC-AGI-3 Semi-Private score from 2.15% to 7.78%, an increase of roughly 3.6 times, while holding the reported Sol variant constant (ARC Prize).
Before drawing any conclusion from that number, there's a second figure on the same page worth mentioning first: ARC Prize also lists a 13.33% Public average for that same Sol configuration at Max effort. That isn't a second confirmation of the 7.78% score. It's a different number from a different evaluation split, and mixing the two up is the most common way this result gets misquoted (ARC Prize).
The stakes go beyond one model's scoreboard. ARC Prize's technical report, published roughly three months ago, describes ARC-AGI-3 as the only unsaturated general agentic intelligence benchmark it's aware of, and its March 2026 data snapshot showed every frontier model tested, Opus 4.6, Gemini 3.1 Pro Preview, GPT-5.4, and Grok-4.20, scoring under 1% (ARC-AGI-3 Technical Report). Sol's Max-effort score arrived about four months after that snapshot, so measured against it, 7.78% is a real jump.
This piece walks through what actually changed in the GPT-5.6 Sol ARC-AGI-3 score, why that change affected Sol far more than the other two GPT-5.6 variants, what ARC-AGI-3 is built to measure, and where a documented result ends and a reasonable guess about causes begins. It's written for high school and college students studying AI, and for teachers looking for a concrete case study in reading a benchmark table without overclaiming what it shows.
The score comparison that matters: Public versus Semi-Private

ARC Prize's results table reports two variables for every score: which variant was used (Sol, Terra, or Luna) and how much reasoning effort that variant was allowed (Low, Medium, High, Extra High, or Max) (ARC Prize). Those are the only two dimensions the table documents. It doesn't show whether a typical user can adjust these values directly or whether they're fixed configurations ARC Prize tested internally.
Within Sol specifically, the ARC-AGI-3 Semi-Private score rises at every step: 0.33% at Low, 1.07% at Medium, 2.15% at High, 6.99% at Extra High, and 7.78% at Max. The table below lays out the same measure across all three variants.
Effort setting Sol Terra Luna Low 0.33% 0.01% 0.17% Medium 1.07% 0.08% 0.17% High 2.15% 0.49% 0.10% Extra High 6.99% 0.65% 0.02% Max 7.78% 0.80% 0.18%
ARC-AGI-3 Semi-Private scores by variant and reasoning-effort setting (ARC Prize).
Notice where most of Sol's gain happens. The jump from High to Extra High (2.15% to 6.99%) accounts for more of the total increase than the jump from Extra High to Max (6.99% to 7.78%). Anyone summarizing this as "Max effort tripled the score" is skipping the fact that most of that gain showed up one setting earlier.
The 13.33% Public figure and the 7.78% Semi-Private figure both describe Sol at Max effort, but they come from different environment pools. The public demo contains 25 environments; ARC-AGI-3 contains 135 environments in total, and the Semi-Private split is the one ARC Prize treats as its official measure for comparing models (ARC Prize; ARC-AGI-3 Technical Report). Sol also became the first model to outright win a public ARC-AGI-3 game, completing an environment called "ft09" at 87% efficiency. That's a genuine milestone, but it comes from a single public environment, not the private-set number used for cross-model comparisons (ARC Prize).
What this means for anyone citing the number in a paper, a lesson, or a discussion post: check whether a reported ARC-AGI-3 score is Public or Semi-Private before comparing it to another model's score, and confirm both numbers hold the variant and effort setting constant. A 13.33% and a 7.78% describing the same model at the same effort level aren't competing claims. They're answers to two different questions.
What the results show across Sol, Terra, and Luna

It's tempting to read Sol's climb and conclude that raising reasoning effort simply works, full stop. The table doesn't support that as a general rule. At Max effort, Terra reaches 0.80% and Luna reaches 0.18%, both far below Sol's 7.78% under the identical effort label, which matters when comparing GPT-5.6 Sol max reasoning effort settings to Terra's or Luna's own Max configurations (ARC Prize).
Terra's climb is real, even if it looks small next to Sol's. Its score moves from 0.01% at Low to 0.80% at Max, an increase of roughly 80 times relative to its own starting point. Luna's pattern is stranger: it opens at 0.17% at Low, holds at 0.17% at Medium, drops to 0.10% at High and 0.02% at Extra High, then rises back to 0.18% at Max. That's not a flat line and it's not a steady climb. Across the ARC-AGI-3 reasoning effort settings tested, Luna's number moves in both directions, and nothing in ARC Prize's published data explains why.
One useful way to separate what's documented from what's speculation: the table shows an observed pattern, not a confirmed cause. As the labeled reasoning-effort setting rises, Sol's reported ARC-AGI-3 score rises with it, while Terra's rises more modestly and Luna's moves inconsistently. A reasonable hypothesis, worth stating as a hypothesis, is that Sol's underlying setup makes better use of additional reasoning steps than Terra's or Luna's do. What's missing is a confirmed mechanism. The supplied ARC Prize sources don't describe what distinguishes Sol's design from Terra's or Luna's, and no replay analysis of Sol's own runs has been published alongside this results page (ARC Prize).
A quick worked example shows why this distinction matters for anyone checking the math. Divide Sol's Max score by its High score: 7.78 divided by 2.15 equals 3.62, which is where the "3.6 times" figure in the opening of this piece comes from. That ratio describes one variant, on one benchmark, between two specific settings. It says nothing about whether Terra shows the same multiple (it doesn't: Terra's High-to-Max ratio is closer to 1.6), whether a different benchmark would show the same pattern, or whether a future model would repeat it. Treating a single ratio as a general rule about "reasoning effort" is the most common way this kind of result gets overstated.
On the older, more saturated ARC-AGI-1 and ARC-AGI-2 benchmarks, effort still helps, but the gains are proportionally smaller. Sol's ARC-AGI-1 score moves from 74.5% at Low to 96.5% at Max, a solid improvement, but nowhere near the multiple seen on ARC-AGI-3 (ARC Prize). That gap between benchmarks is itself informative: it suggests ARC-AGI-3's novelty and lack of instructions give reasoning effort more room to matter than a benchmark where models have more practice-relevant patterns to lean on.
What the ARC-AGI-3 benchmark analysis actually measures

Understanding why a 7.78% score counts as meaningful requires understanding what the test is built to do. ARC-AGI-3 consists of 135 hand-built environments where an agent receives no instructions at all. It has to explore a 64x64 grid, figure out what the hidden goal even is, and plan a sequence of actions from a limited set of controls (ARC-AGI-3 Technical Report).
Every environment went through a validation process built to rule out lucky guessing. A random, uninformed policy couldn't solve any given level more than once in 10,000 attempts, so a decent score can't come from an agent stumbling into the answer (ARC-AGI-3 Technical Report).
The scoring method is where a lot of confusion starts. A 100% AI score doesn't mean an AI system solved every level outright. It means the system matched or exceeded the median human action-efficiency baseline across every private environment, a bar set by how efficiently ordinary people solve the same unfamiliar tasks (ARC-AGI-3 Technical Report). ARC Prize's separate human-performance study, released roughly three and a half months ago, backs this up with numbers: every environment in the study was solved by at least two participants, and typically five or more, out of a panel of around ten people with no special training (ARC Prize human-dataset post).
Set against that baseline, Sol's 7.78% Semi-Private score sits far below the point where it would match how efficiently people solve the same tasks. It's a genuine advance compared with a field of frontier models scoring under 1% in that March 2026 snapshot, but "genuine advance" and "closing in on human-level performance" are two different claims, and only the first one is supported by these numbers. Anyone citing the "below 1%" comparison later should note that it's tied to that specific snapshot, not a current standing, since it predates the GPT-5.6 Sol results by months.
What earlier replay research can and can't explain about Sol's gain

There's one more piece of evidence worth including, with a clear boundary attached. ARC Prize published a replay analysis, roughly three months ago, examining how GPT-5.5 and Opus 4.7 behaved on ARC-AGI-3 environments, action by action, alongside their reasoning traces. That analysis does not include GPT-5.6 Sol. No behavior-level replay of Sol's Max-effort runs has been published, so what follows explains a plausible mechanism drawn from other models, not a confirmed account of what happened inside Sol's own runs (ARC Prize replay analysis).
In that study, both models could sometimes notice a local effect, like recognizing that pressing a particular action rotates an object on screen, without ever turning that observation into a workable model of how the environment worked as a whole. A second failure mode showed up repeatedly too: models explained brand-new game mechanics by mapping them onto familiar patterns from training data, like Tetris, Sokoban, Frogger, or Pong, even when the actual rules didn't match (ARC Prize replay analysis).
One documented run makes the risk concrete. Opus solved the first level of an environment called "ka59" using an incorrect theory about how a mechanic worked. When the second level required the real mechanic, Opus's wrong theory had already hardened into a fixed rule, and the run never recovered (ARC Prize replay analysis).
It's reasonable to infer that extra reasoning effort gives a model more chances to test a hypothesis and revise it before locking in, which would help explain why effort tracks so strongly with score on this particular benchmark. That inference is plausible and consistent with the mechanism documented in these other models. It's still an inference, though, not a description of what Sol's own reasoning traces show, since none have been published.
For a classroom exercise that uses this distinction directly: have students line up Sol's Low, High, Extra High, and Max scores, then write two separate sentences about the jump from High to Max. One sentence should state only what the table shows. The second, clearly labeled as an inference, should propose a reason for it. Keeping those two sentences separate is the core skill behind reading any AI benchmark claim without overclaiming it.
Checking a benchmark claim before using it
Before citing the GPT-5.6 Sol ARC-AGI-3 score in coursework, a lesson plan, or a conversation about AI progress, run through a short checklist:
- Record the benchmark version and the date the score was published.
- Note the model, the variant, and the exact reasoning-effort or compute setting.
- Confirm whether the number is a Public or Semi-Private score.
- Check what the scoring method actually measures, not just what the percentage implies.
- Look for replay or behavioral evidence before accepting an explanation of why a score changed.
That checklist also doubles as a starting point for anyone researching how to improve ARC-AGI-3 benchmark scores: settings alone aren't the answer, since the same effort label produced a steady climb for Sol, a modest one for Terra, and an inconsistent one for Luna. Applying the checklist to this result gives a specific, defensible summary: within the Sol variant, the reported ARC-AGI-3 Semi-Private score rose from 2.15% to 7.78% as reasoning effort increased from High to Max, a documented association confirmed on ARC Prize's own results page. Terra and Luna, tested under the same effort labels, didn't show the same pattern, and no published source explains why.
Before repeating a stronger version of this claim in a paper or class discussion, open ARC Prize's results page and technical report directly, confirm which split and settings a given number reflects, and check whether a replay analysis of Sol's own runs has since been published.