
A NVIDIA-built coding agent called AVO cleared every one of the 183 levels across 25 public games on the ARC-AGI-3 benchmark, finishing at 100%. The same underlying model, Anthropic’s Claude Opus 5 running without the harness, scored 30% on the identical set, a gap produced by tooling rather than by any change to the model itself.
What AVO is and what it did
AVO is a software harness, a wrapper that exposes tools and a feedback loop to a language model, wrapped around Claude Opus 5. NVIDIA did not change the core agent architecture. It swapped the GPU engineering tools that AVO was originally built around for the ARC-AGI-3 task interface, then let the agent run.
On the ARC-AGI-3 public set the result was a perfect score: all 183 levels across 25 public games. The agent received no rules, no prior instruction, and no stated goals. It learned by trying actions, observing the results, and correcting itself.
Why the same model behaves so differently
Claude Opus 5 on its own, without the AVO harness, scored 30% on the same public set. The 70-point gap between 30% and 100% comes from the wrapper: the loop that lets the model probe the environment, see what changed, and try again. Swap the harness and the same underlying model moves from a failing score to a clean sweep.
Where AVO came from
AVO was originally built to optimise CUDA GPU kernels. In that earlier role it ran autonomously for 7 days, explored more than 500 design directions, and produced kernels that beat FlashAttention-4 by up to 10.5%. The same agent, redirected at a completely different problem class, picked up ARC-AGI-3 without retraining.
How efficient was the run?
AVO cleared the 183 levels in 6,624 actions. That is roughly 12% more efficient than VISTA, which needed 7,542 actions on the same public set. Fewer actions for a perfect score points to a tighter search, not just a longer one.
What is not yet known
ARC-AGI-3 does not allow external harnesses to run against its hidden private set, so AVO’s private-set performance is unknown. The 100% figure and the efficiency comparison both apply to the public set only.
Why the harness matters
The headline lesson is that a model’s score on a benchmark is a property of the model plus its wrapper, not of the model alone. The same Claude Opus 5 that scored 30% on its own reached 100% inside AVO, with no model-side changes. For teams evaluating agent benchmarks, the harness, the tools it exposes, and the feedback loop it closes are part of the result.
FAQ
What is AVO?
AVO is NVIDIA’s coding agent, a software harness wrapped around Anthropic’s Claude Opus 5. It was originally built to optimise CUDA GPU kernels and was redirected at ARC-AGI-3 by swapping its tools.
What did AVO score on ARC-AGI-3?
AVO scored 100% on the ARC-AGI-3 public set, clearing all 183 levels across 25 public games. Claude Opus 5 without the harness scored 30% on the same set.
Why does the private-set score matter?
ARC-AGI-3 does not allow external harnesses to run against its hidden private set, so AVO’s private-set performance is unknown. The 100% result applies only to the public set.
See our editorial disclaimer for how our articles are produced.
Run a free scan to see your AI Visibility Score, SEO rating, and local citation accuracy.