When Agent Governance Helps
Michael Ray Johnson, Linda Naimi
Submitted to arXiv on September 2, 2026
Abstract
As autonomous AI agents increasingly operate within defined guardrails, a fundamental question arises: how should a governed autotelic multi-agent organization—where agents pursue self-generated goals under constraints—be designed and evaluated? This paper addresses that question in two parts.
First, we present the Governed Autotelic Multi-Agent Product Organization (GAMPO) framework, developed from a document-based qualitative evidence synthesis of 321 sources. GAMPO integrates agency theory, agile methodologies, platform design, and governance theory into a runnable specification. Second, we empirically probe a prompt-layer instantiation of GAMPO on CHI-Bench, a long-horizon healthcare benchmark, across both open and frontier models.
Our findings establish a critical boundary condition: the benefit of governance is gated by a model's spare capacity and is both domain-specific and model-specific. On capacity-constrained open models, the full procedure yields no reliable improvement. However, a minimal intervention—a single sentence instructing the agent to 'verify your writes'—doubles task success (pass@1 improving from 2/20 to 4/20). At the frontier, the same scaffold lifts prior-authorization performance from 24% to 40% on one model but yields zero benefit on another, a gap traced to a stable recommendation-override disposition.
A second result refines this picture. When the generic procedure is replaced with an answer-blind, per-task definition-of-done—keyed only to the case's own policy and published standards, never the hidden ground truth—performance jumps to 84% on prior-authorization under best-of-five self-consistency (with 68% single-attempt success, confirmed by a held-out board). Utilization-management reaches 44%, while care-management encounters a ceiling on subjective content quality.
The contribution is twofold: a named, auditable framework, and capability-gated evidence showing that governance should be sized to available capacity. At the frontier, a case-grounded specification outperforms a uniform procedure. These findings are exploratory, based on partial instantiation, small per-cell samples (n = 5–25), and single trials.
Keywords: autotelic AI agents, multi-agent governance, GAMPO framework, CHI-Bench, healthcare AI, capability-gating, protocol design, self-governance
Note: This is a companion empirical study to a Purdue D.Tech dissertation. The paper introduces the GAMPO framework from a 321-source qualitative evidence synthesis and probes its prompt-layer instantiation on the CHI-Bench healthcare benchmark.
via ArXiv CL+LG
