Nvidia published new research on Friday highlighting that the harness, not the AI model itself, is critical for enabling long-horizon tasks. A harness is the software framework that wraps around an AI model, providing tools, memory management, and rules to make it function independently. According to the study, a custom harness tailored for memory management and including a 'supervisor' component allowed Claude Opus 5 to achieve a perfect 100% score on the interactive reasoning benchmark ARC-AGI-3. This benchmark, which consists of 2D games with no instructions, requires models to figure out how to play and win, much like a human would. Without the harness, Opus 5 scored only 30%, which was the top result among all tested models. This demonstrates that while model choice matters, the harness plays a more significant role in enabling agentic behavior, especially for complex, long-term tasks.

The research underscores the importance of the harness in managing memory, context, and feedback, which are essential for an AI to perform effectively over extended periods. Adel El Hallack, vice president of product at Nvidia’s AI unit, explained that the harness acts as the scaffolding around the model, providing the runtime and skills needed for the model to operate as an agent. He noted that the harness is more than just an API for the model, as it includes the tools, libraries, and infrastructure that enable the model to function autonomously. This is particularly important for long-horizon tasks, which require stringing together many decisions over time, often spanning days. Such tasks are far more complex than simple prompt-based responses, and the challenge lies in ensuring the AI stays focused and does not get distracted.

Nvidia’s findings add to a growing body of evidence that model choice is not the only factor in agentic performance. In July, Databricks published research showing that the harness has a more significant impact on AI costs than the model itself. Ali Ghodsi, Databricks CEO, explained that using the wrong harness could double the cost of running a model, regardless of its inherent efficiency. Nvidia’s larger point is to emphasize that open harnesses, like open models, give users more control over AI systems. El Hallack argued that open harnesses allow for greater customization, enabling users to fine-tune performance and accuracy. This aligns with Nvidia’s broader vision of fostering an open agent stack that gives users control across the harness, infrastructure, and runtime, ensuring secure and efficient AI development.

Source: techcrunch