{"id":399338,"date":"2026-08-23T17:30:14","date_gmt":"2026-08-23T17:30:14","guid":{"rendered":"https:\/\/bizscoreai.com\/blog\/avo-beats-claude-opus-5-arc-agi-3-harness\/"},"modified":"2026-08-23T17:30:16","modified_gmt":"2026-08-23T17:30:16","slug":"avo-beats-claude-opus-5-arc-agi-3-harness","status":"publish","type":"post","link":"https:\/\/bizscoreai.com\/blog\/avo-beats-claude-opus-5-arc-agi-3-harness\/","title":{"rendered":"AVO Beats Claude Opus 5 on ARC-AGI-3 by Changing the Harness, Not the Model"},"content":{"rendered":"<p>A NVIDIA-built coding agent called AVO cleared every one of the 183 levels across 25 public games on the ARC-AGI-3 benchmark, finishing at 100%. The same underlying model, Anthropic&#8217;s Claude Opus 5 running without the harness, scored 30% on the identical set, a gap produced by tooling rather than by any change to the model itself.<\/p>\n<h2>What AVO is and what it did<\/h2>\n<p>AVO is a software harness, a wrapper that exposes tools and a feedback loop to a language model, wrapped around Claude Opus 5. NVIDIA did not change the core agent architecture. It swapped the GPU engineering tools that AVO was originally built around for the ARC-AGI-3 task interface, then let the agent run.<\/p>\n<p>On the ARC-AGI-3 public set the result was a perfect score: all 183 levels across 25 public games. The agent received no rules, no prior instruction, and no stated goals. It learned by trying actions, observing the results, and correcting itself.<\/p>\n<h2>Why the same model behaves so differently<\/h2>\n<p>Claude Opus 5 on its own, without the AVO harness, scored 30% on the same public set. The 70-point gap between 30% and 100% comes from the wrapper: the loop that lets the model probe the environment, see what changed, and try again. Swap the harness and the same underlying model moves from a failing score to a clean sweep.<\/p>\n<h2>Where AVO came from<\/h2>\n<p>AVO was originally built to optimise CUDA GPU kernels. In that earlier role it ran autonomously for 7 days, explored more than 500 design directions, and produced kernels that beat FlashAttention-4 by up to 10.5%. The same agent, redirected at a completely different problem class, picked up ARC-AGI-3 without retraining.<\/p>\n<h2>How efficient was the run?<\/h2>\n<p>AVO cleared the 183 levels in 6,624 actions. That is roughly 12% more efficient than VISTA, which needed 7,542 actions on the same public set. Fewer actions for a perfect score points to a tighter search, not just a longer one.<\/p>\n<h2>What is not yet known<\/h2>\n<p>ARC-AGI-3 does not allow external harnesses to run against its hidden private set, so AVO&#8217;s private-set performance is unknown. The 100% figure and the efficiency comparison both apply to the public set only.<\/p>\n<h2>Why the harness matters<\/h2>\n<p>The headline lesson is that a model&#8217;s score on a benchmark is a property of the model plus its wrapper, not of the model alone. The same Claude Opus 5 that scored 30% on its own reached 100% inside AVO, with no model-side changes. For teams evaluating agent benchmarks, the harness, the tools it exposes, and the feedback loop it closes are part of the result.<\/p>\n<h2>FAQ<\/h2>\n<h3>What is AVO?<\/h3>\n<p>AVO is NVIDIA&#8217;s coding agent, a software harness wrapped around Anthropic&#8217;s Claude Opus 5. It was originally built to optimise CUDA GPU kernels and was redirected at ARC-AGI-3 by swapping its tools.<\/p>\n<h3>What did AVO score on ARC-AGI-3?<\/h3>\n<p>AVO scored 100% on the ARC-AGI-3 public set, clearing all 183 levels across 25 public games. Claude Opus 5 without the harness scored 30% on the same set.<\/p>\n<h3>Why does the private-set score matter?<\/h3>\n<p>ARC-AGI-3 does not allow external harnesses to run against its hidden private set, so AVO&#8217;s private-set performance is unknown. The 100% result applies only to the public set.<\/p>\n<p><script type=\"application\/ld+json\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"FAQPage\",\"mainEntity\":[{\"@type\":\"Question\",\"name\":\"What is AVO?\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"AVO is NVIDIA's coding agent, a software harness wrapped around Anthropic's Claude Opus 5. It was originally built to optimise CUDA GPU kernels and was redirected at ARC-AGI-3 by swapping its tools.\"}},{\"@type\":\"Question\",\"name\":\"What did AVO score on ARC-AGI-3?\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"AVO scored 100% on the ARC-AGI-3 public set, clearing all 183 levels across 25 public games. Claude Opus 5 without the harness scored 30% on the same set.\"}},{\"@type\":\"Question\",\"name\":\"Why does the private-set score matter?\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"ARC-AGI-3 does not allow external harnesses to run against its hidden private set, so AVO's private-set performance is unknown. The 100% result applies only to the public set.\"}}]}]}<\/script><\/p>\n<hr style=\"margin:2.5em 0 1em;opacity:.35\" \/>\n<p style=\"font-size:.85em;opacity:.7\">See our <a href=\"https:\/\/bizscoreai.com\/blog\/disclaimer\/\">editorial disclaimer<\/a> for how our articles are produced.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>A NVIDIA coding agent wrapped around Claude Opus 5 went from 30% to 100% on ARC-AGI-3 by swapping tools, not the underlying model.<\/p>\n","protected":false},"author":1,"featured_media":399337,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"rank_math_title":"AVO Beats Claude Opus 5 on ARC-AGI-3 via Harness","rank_math_description":"NVIDIA's AVO harness wrapped around Claude Opus 5 scored 100% on ARC-AGI-3 vs 30% for the bare model, using the same architecture.","rank_math_focus_keyword":"avo claude opus 5","footnotes":""},"categories":[1],"tags":[],"class_list":["post-399338","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-news"],"elementor_data":null,"elementor_edit_mode":null,"_links":{"self":[{"href":"https:\/\/bizscoreai.com\/blog\/wp-json\/wp\/v2\/posts\/399338","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/bizscoreai.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/bizscoreai.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"replies":[{"embeddable":true,"href":"https:\/\/bizscoreai.com\/blog\/wp-json\/wp\/v2\/comments?post=399338"}],"version-history":[{"count":1,"href":"https:\/\/bizscoreai.com\/blog\/wp-json\/wp\/v2\/posts\/399338\/revisions"}],"predecessor-version":[{"id":399339,"href":"https:\/\/bizscoreai.com\/blog\/wp-json\/wp\/v2\/posts\/399338\/revisions\/399339"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/bizscoreai.com\/blog\/wp-json\/wp\/v2\/media\/399337"}],"wp:attachment":[{"href":"https:\/\/bizscoreai.com\/blog\/wp-json\/wp\/v2\/media?parent=399338"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/bizscoreai.com\/blog\/wp-json\/wp\/v2\/categories?post=399338"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/bizscoreai.com\/blog\/wp-json\/wp\/v2\/tags?post=399338"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}