Google researchers introduced RRSI, a technique designed to stop self-improving AI agents from memorizing their test tasks, according to a new research paper. The method aims to prevent agents from overspecializing on test tasks while reducing compute costs.

The paper highlights that AI agents using self-optimization often memorize test tasks, leading to poor performance on new, unseen benchmarks. RRSI addresses this by introducing constraints that limit how much the harness can change during self-optimization, preventing the agent from hardcoding solutions or task-specific tricks.

RRSI works by shrinking an edit budget over time, allowing larger rewrites in early rounds and smaller, traceable changes in later rounds. It also maintains a record of previous attempts to avoid repeating failed strategies. When progress stalls, the system deliberately experiments with parts of the harness it hasn't touched yet.

"The system produces feedback that it uses to optimize the harness, which in turn controls the system's own behavior," said the researchers. This feedback loop ensures the harness remains generalizable while still allowing for self-improvement.

The researchers tested RRSI on eight benchmarks, including coding, office work, and engineering design. The underlying model, Claude Opus 4.8, stayed frozen throughout the experiments. RRS, which stands for Regularized Recursive Self-Improvement of Agent Harnesses, outperformed other methods on unseen tasks, achieving up to 4.7-point gains.

Google researchers noted that RRSI requires fewer tokens and steps than other methods and performs best on new tasks. However, they emphasized that their study only covers harnesses built around frozen models and does not address cases where the model weights change. The code is available on GitHub.

Source: thedecoder