Heavy-Tailed Memory Traces in Long-Horizon Language Agents

Heavy-Tailed Memory Traces in Long-Horizon Language Agents


arXiv:2610.00010 [cs.AI] | Submitted 9 July 2026 | Under Review


Authors: Xinyuan Song, Zekun Cai


Abstract


Long-horizon language agents increasingly rely on external memory as a frozen world model, yet current memory systems are usually judged only by task success or token cost. We argue that the missing object is the shape of memory use: under finite context and repeated retrieval, agent memory can concentrate on a small core while leaving rare states in a long tail where prediction errors accumulate.


We study this effect through a conservative tail audit and find that concentration is reproducible but policy-dependent. Random-walk agents produce log-normal-compatible retrieval artifacts, whereas semantic LLM policies yield the strongest truncated-power-law-compatible core–tail traces.


Motivated by this audit, we propose Core–Tail World Model (CTWM), a rank-based memory controller that allocates prompt budget with a single exponent $\tau$ while retaining a summarized tail. On Synthetic Graph World, CTWM:


  • Preserves full state and transition coverage
  • Reduces prompt tokens by 5.9%
  • Lowers bottom-half tail prediction error by 13.6% relative to a graph-memory baseline

The same paired comparison gives consistent token savings on ALFWorld and a 24.48% token reduction on LongMemEval with aggregate accuracy parity.


These results suggest that heavy-tailed memory traces are not only a diagnostic of finite retrieval, but also a practical control signal for token-efficient agent world models.


Key Contributions


  1. Diagnostic framing: Identifies the shape of memory use—rather than task success or token cost alone—as a critical, previously overlooked property of long-horizon language agents.
  2. Tail audit: Demonstrates that memory concentration is reproducible but driven by agent policy. Random-walk agents yield log-normal-compatible artifacts, while semantic LLM policies produce truncated-power-law-compatible core–tail traces.
  3. CTWM controller: Introduces a rank-based memory controller that allocates prompt budget via a single exponent $\tau$, retaining a summarized tail to guard against accumulated prediction errors.
  4. Empirical validation: Achieves token savings and reduced tail prediction error across Synthetic Graph World, ALFWorld, and LongMemEval without sacrificing aggregate accuracy.

  5. Metadata


    • Comments: Under Review
    • Subjects: Artificial Intelligence (cs.AI)
    • Cite as: arXiv:2610.00010 [cs.AI]
    • DOI: https://doi.org/10.48550/arXiv.2610.00010
    • Submission history: [v1] Thu, 9 Jul 2026 15:00:39 UTC (3,093 KB)

    2026 Context


    As language agents move toward longer horizons and more persistent external memory in 2026, efficiency and reliability at scale become central design constraints. Prior work has largely optimized memory for task success or raw token cost, but the distributional shape of memory access—how retrieval concentrates on a small core versus a rare-state tail—has remained underexplored. This paper positions heavy-tailed memory traces as both a diagnostic signal and a practical control handle, offering a lightweight, rank-based controller (CTWM) that trades a single tunable exponent for measurable token and error gains across standard agent benchmarks.

    via ArXiv AI

Related