
ARC-AGI-3's perfect score belongs to the harness
Nvidia's AVO scored 100.00 on ARC-AGI-3's public set while ARC Prize reports Claude Opus 5 High at 30.16%—but Nvidia says this does not isolate harness gain.
ARC Prize reported Claude Opus 5 High at 30.16% on ARC-AGI-3. Nvidia then reported 100.00 across the benchmark's 25-environment public set using the same model family inside its AVO agent system.
The tempting subtraction is seventy points of harness gain. Nvidia explicitly says not to make it. The two results use the same model family under a different reasoning setting and a substantially different agent system and evaluation setup. The result shows that a model name no longer identifies the system being benchmarked; it does not isolate how many points the harness caused.
That narrower claim is still enough to break the ordinary model bake-off.
Nvidia reported a complete-system result
Nvidia's August 21 technical report describes AVO as a long-horizon coding-agent architecture. Its main agent cycles through inspect, plan, implement and evaluate. Persistent memory carries prior implementations and test results forward. A supervisor watches the broader trajectory and redirects the main agent when progress stalls.
For ARC-AGI-3, Nvidia connected that architecture to a text-only interface. Each observation arrived as an exact 64 × 64 text grid; the agent received available actions but no description of the game's rules or goal. AVO completed all 183 public levels and received a 100.00 Relative Human Action Efficiency score.
That score is not completion alone. ARC Prize's scoring method combines levels completed with environment-action efficiency against first-time human baselines. A 100% total means every level was completed while matching or exceeding the human-efficiency bar.
So the achievement is real and specific: Nvidia reported one full agent system clearing the public set at that bar. It is not evidence that AVO clears the semi-private or private competition sets, and Nvidia makes no such claim.
The comparison changes more than the harness
The model reference is also real. ARC Prize's July 24 result reports Claude Opus 5 High at 30.16%, the best model score at the time. But Nvidia says its 100-point run used the same model family under a different reasoning setting and a substantially different system and evaluation setup.
“Harness multiple” is too neat a label for this comparison. A ratio needs a controlled numerator and denominator; these two rows change too much at once. Calling the observed gap a seven-to-one harness effect would turn Nvidia's warning into decoration.
Fewer environment actions do not establish cheaper, faster or safer operation. The cited metric does not include latency, internal reasoning steps or cost.
The defensible inference is more useful for procurement anyway: when a familiar model family sits inside a system with its own observation interface, memory and supervision, then “we tested Opus” is not a reproducible evaluation record. It is the start of one.
A ceiling changes the next measurement
The old article called the perfect score an obituary for the benchmark. That was too broad. ARC-AGI-3 still has semi-private and private sets, and the public set remains useful for debugging and comparison below the ceiling.
What 100.00 ends is the public score column's ability to distinguish AVO from any future system that also reaches 100. Nvidia gives one useful second measurement: AVO used 6,624 environment actions, against VISTA's 7,542 while both completed all 183 public levels—about 12% fewer.
Even that is a cross-system comparison, not a controlled ablation. And environment actions are not wall-clock time, inference cost or internal reasoning steps. Fewer game actions may reflect a better policy without producing a cheaper deployment. Report all four separately instead of asking one benchmark number to stand in for them.
Version the system you actually tested
For the next agent bake-off, keep each run as one versioned system row: exact model and reasoning setting; observation representation and tools; memory and context policy; supervisor, retry and stop rules; environment budget; then completion, RHAE, environment actions, wall-clock time, model and tool cost, human interventions and reversals.
Change one row at a time when you want causality. Change the whole system when you want the best production result. Do not run the second experiment and describe it as the first.
This advances an earlier archive finding that model leaderboards cannot see the harness. The earlier piece shows the same model moving across two runtimes. This result adds the ceiling case: once a complete system saturates a public score, the evaluation has to expose the system version and move to harder evidence. It also sharpens the release-platform rule: the whole agent, not just the model, is the artifact to test and reverse.
Nvidia's result does not make models irrelevant. It makes an unlabeled model column incomplete. Preserve the 100.00, attach the public-set scope, and version every surrounding choice that helped earn it. Then the next team can reproduce the system instead of repeating the headline.
A perfect agent score belongs to the versioned system that earned it—benchmark the model, harness and evaluation setup together, then isolate variables before claiming why the score moved.


