Agent-Controlled Forgetting: Reversible Context Curation for

Agent-Controlled Forgetting: Reversible Context Curation for Tool-Using AI Agents


arXiv:2610.10590 (cs.AI) | Submitted 6 Oct 2026 | 13 pages, 1 figure


Author: Jan-Peter Franke


Code and research artifacts: github.com/boldprojekte/agent-forgetting-research




Abstract


Tool-using agents repeatedly carry observations whose useful content can be much smaller than their original payload. We study agent-controlled forgetting: the acting model selects previously observed tool results, replaces each with a short note at its original position, and retains the exact original in a recoverable archive. A Python harness exposes batch archival and explicit recovery without task-specific model training, while protecting user instructions and assistant messages from these operations.


In an exploratory OpenTelemetry debugging case followed by an unrelated implementation task, the method ended with 231,951 provider-reported prompt tokens versus 912,492 under retained history, used 50% fewer cumulative input tokens, and had an estimated API cost of USD 1.28–1.44 versus approximately USD 4.38. Both arms passed the two-case primary behavioral oracle; neither fully satisfied the follow-up evaluation. The method made more requests and took 17% longer.


A contrasting application-development pair produced no context or cost saving, and an earlier continuation exhibited lower manually assessed quality despite reduced context. These observations demonstrate substantial resource savings in noisy tool-use trajectories and identify workload dependence as a central consideration for reversible context management.




Key Contributions


  • Agent-controlled forgetting mechanism: The acting model itself decides which previously observed tool results to archive, replacing each with a concise note at its original position while preserving the exact original in a recoverable archive.
  • Reversible context curation: All forgetting operations are recoverable β€” nothing is permanently lost, and the archive can be queried for exact originals on demand.
  • Training-free implementation: A lightweight Python harness enables batch archival and explicit recovery without task-specific model fine-tuning.
  • Protected context: User instructions and assistant messages are shielded from archival operations, preserving alignment and conversational integrity.



Experimental Results


| Metric | Agent-Controlled Forgetting | Retained History |

|---|---|---|

| Final prompt tokens (provider-reported) | 231,951 | 912,492 |

| Cumulative input tokens | 50% reduction | Baseline |

| Estimated API cost | USD 1.28–1.44 | ~USD 4.38 |

| Latency | +17% | Baseline |

| Primary behavioral oracle | Passed | Passed |

| Follow-up evaluation | Partially satisfied | Partially satisfied |


Notable Findings


  • Substantial token and cost savings were observed in noisy tool-use trajectories β€” the OpenTelemetry debugging case being the clearest example.
  • The method required more requests and 17% longer wall-clock time, trading latency for context economy.
  • In a contrasting application-development pair, no context or cost saving materialized, underscoring that benefits are workload-dependent.
  • An earlier continuation showed lower manually assessed quality despite reduced context, suggesting a quality–efficiency trade-off worth monitoring.



Implications for 2026 Agent Design


As agentic systems in 2026 routinely chain dozens of tool calls per session β€” browsing, code execution, database queries, and file operations β€” the context window has become both a technical and an economic bottleneck. Prompt tokens remain the primary cost driver for most frontier models, and unbounded observation accumulation degrades both latency and reasoning quality.


Agent-controlled forgetting offers a middle path between naive truncation and full retention:


  • Versus truncation: Truncation is destructive and irreversible; reversible curation permits recovery when downstream reasoning demands the original payload.
  • Versus slim summarization: Summarization discards fidelity; archival notes preserve a path back to exact originals.
  • Versus external memory injection: No retrieval infrastructure is required at inference time β€” the mechanism operates entirely within the agent's own control loop.

Workload dependence is the critical caveat. The paper shows that for noisy, high-volume tool trajectories the savings are dramatic, while for lean application-development workflows the overhead may yield no benefit. Practitioners should profile their own agent workloads before adopting reversible context curation as a default strategy.




Citation


@article{franke2026agentcontrolled,
  title={Agent-Controlled Forgetting for Tool-Using Agents: Reversible Context Curation in Practice},
  author={Franke, Jan-Peter},
  journal={arXiv preprint arXiv:2610.10590},
  year={2026},
  doi={10.48550/arXiv.2610.10590}
}

Subjects: Artificial Intelligence (cs.AI)


Cite as: arXiv:2610.10590 [cs.AI]


DOI: 10.48550/arXiv.2610.10590

via ArXiv AI

Related