Google Research Open-Sources RRSI: AI Agents That Improve Their Own Harness Without Overfitting
Google Cloud AI Research, in collaboration with UNC-Chapel Hill, Stanford, and Washington University in St. Louis, has released RRSI (Regularized Recursive Self-Improvement). The framework enables an LLM agent to rewrite its own harness—including prompts, tools, memory, control flow, and sub-agents—without ever changing the underlying model weights. RRSI constrains the improvement loop itself, ensuring that gains generalize to benchmarks the agent never optimized against.
Deployable as a Research Framework
RRSI is available as a research framework under the Apache 2.0 license. It requires Python 3.10+ and accepts any LiteLLM model string, with defaults assuming Claude Opus 4.8 on Vertex AI.
Why Self-Improving Harnesses Overfit
Harness evolution loops typically propose edits, score them on a fixed evolve set, and keep the winner. Because the same tasks are reused every round, the loop can memorize them. The RRSI research identifies three failure modes:
- Benchmark-specific fitting
- Noise chasing
- Complexity accumulation
- Annealed edit budget: A cosine schedule allows early rounds to bundle several edits, while late rounds permit only a single attributable change.
- Evidence-aware credit: Each candidate is logged with its supporting evidence, so improvements can be traced to specific modifications.
Each one widens the gap between evolve-set scores and real-world transfer.
How RRSI Works
RRSI keeps every harness component editable. Instead of restricting what can change, it regularizes how the search moves.
Proposal Side
via MarkTechPost
