Synthesis Through Simulation: Generating Coherent Enterprise

Synthesis Through Simulation: Generating Coherent Enterprise Data via Scalable Agent-System Interaction


Authors: Yipeng Li, Ashutosh Hathidara, Jane Lo, Harshavardhan Abichandani, Gunraj Singh, Atin Ghosh

Submitted: 24 September 2026

arXiv: 2610.10549 [cs.AI]

DOI: 10.48550/arXiv.2610.10549


Abstract


Tool-calling agents have become central to enterprise AI, yet training and evaluating them at scale remains severely constrained by business and legal restrictions on enterprise systems, data, and database schemas. Tabular data synthesis offers a natural alternative, but its effectiveness is fundamentally limited by structural validity and schema availability. In contrast, procedure-based approaches exhibit the opposite weakness—typically lacking distributional fidelity without per-domain authoring.


We introduce Synthesis Through Simulation (STS), a schema-free data synthesis paradigm in which an LLM agent generates data by executing operations against policy-enforcing APIs within simulated enterprise environments. Because data is generated through the same environment that defines what is valid, STS guarantees structural validity by construction while decoupling validity enforcement from distribution modeling, allowing each to be addressed independently.


The Generalist Populator (GP), STS's domain-agnostic agent, addresses the remaining challenges of distributional fidelity and synthesis scalability. GP achieves 0.88 average marginal fidelity and 100% constraint satisfaction across all ten environments without access to DB schemas. By comparison, statistical synthesizers are inapplicable to seven environments due to necessary seed data requirements, and schema-privileged agents fail 82% of trajectories in the airline environment's tightly coupled workflows due to brittle task composition.


The full framework, all ten environments, and generated datasets are open-sourced at https://github.com/SAP/synthesis-through-simulation.


Key Contributions


  • Schema-free synthesis paradigm: STS generates valid enterprise data without requiring database schemas, removing a major barrier to scalable agent training and evaluation.
  • Validity by construction: By simulating the policy-enforcing environment itself, STS ensures structural validity without separate validation passes.
  • Decoupled validity and distribution: Validity enforcement and distribution modeling are handled independently, enabling each to be optimized separately.
  • Generalist Populator (GP): A domain-agnostic agent that generalizes across ten enterprise environments, achieving high fidelity and full constraint satisfaction.
  • Open-source release: All environments, the framework, and generated datasets are publicly available.

Why It Matters in 2026


As enterprise AI adoption accelerates in 2026, organizations face mounting pressure to train and evaluate tool-calling agents on realistic data—yet privacy regulations, proprietary schemas, and legal constraints continue to block access to production systems. STS offers a practical path forward: enterprises can generate coherent, policy-compliant synthetic data at scale without exposing sensitive infrastructure or requiring per-domain engineering.


This approach is particularly relevant as agentic workflows move from pilot programs into mission-critical deployments, where robustness across tightly coupled tasks—such as those seen in airline, finance, and healthcare systems—is essential.


Resources


via ArXiv AI

Related