When Do Causal World Models Help Modular LLM Agents

When Do Causal World Models Help Modular LLM Agents


Authors: Xinyuan Song, Zekun Cai

arXiv: 2610.00012 [cs.AI]

Submitted: 12 July 2026

Status: Under Review

DOI: https://doi.org/10.48550/arXiv.2610.00012


Abstract


LLM agents increasingly operate through modular systems—such as order, payment, inventory, and shipment services—where actions in one module change which transitions are valid in another. Standard world models typically fit observational traces, but this objective does not capture the quantity needed for intervention-time planning. A trace may show that payment precedes shipment without revealing whether payment authorizes shipment, whether inventory mediates the effect, or whether a hidden trigger explains both.


This paper studies that gap through FedCausalCompose, a causal world-model framework for modular LLM agents in which local actions provide intervention-response evidence for cross-module interfaces.


The authors establish three theoretical results:


  1. Observational world models incur an irreducible interventional error under unblocked back-door paths.
  2. Interface recovery improves as intervention-response coverage grows.
  3. An oracle causal composition can beat the non-causal lower bound when coverage and local mechanism errors are controlled.

  4. These predictions are then tested in diagnostic agent settings. Causal interfaces help most in structured tool environments, where API signatures expose preconditions and downstream effects. By contrast, dialogue and narrative environments often ignore raw edge lists unless a short attention anchor renders the causal information decision-relevant.


    Key Findings


    The results identify a concrete condition for causal world models in LLM agents: causal structure helps when cross-module interfaces are both statistically identifiable and presented in a form the agent can use at action time.


    Why This Matters in 2026


    As LLM agents move from single-turn prompting toward multi-service orchestration—spanning payment gateways, logistics APIs, and enterprise resource systems—the reliability of interventional reasoning becomes a first-order engineering concern. FedCausalCompose offers a principled way to decide when investing in explicit causal structure pays off, and when observational world models remain sufficient.


    Subjects


    • Artificial Intelligence (cs.AI)

    Cite As


    Xinyuan Song, Zekun Cai. When Do Causal World Models Help Modular LLM Agents.
    arXiv:2610.00012 [cs.AI], 2026. https://doi.org/10.48550/arXiv.2610.00012
    



    Submission history available via arXiv. From: Zekun Cai.

    via ArXiv AI

Related