Recognition, Simulation, and Refusal: A Contamination-Aware

Recognition, Simulation, and Refusal: A Contamination-Aware Study of Classic Psychological Effects in LLM Agents


arXiv:2609.22090 [cs.CL] | Submitted 23 Jul 2026 | 13 pages, 1 figure, 8 tables


Author: Joy Bose


Subjects: Computation and Language (cs.CL) | ACM Classes: I.2.7; I.2.11; J.4




Abstract


An LLM producing the response pattern associated with a human psychological effect is not the same claim as the LLM possessing that bias. We present PsyAgentBench, a benchmark that re-runs classic psychology experiments on LLM agents under a factorial design built to separate these possibilities: each paradigm is run with the paradigm explicitly labeled in the prompt (named) or framed as a routine task (blind), and on the literal textbook version of the task (canonical) or a structurally matched variant written to reduce lexical and scenario overlap with likely training data (counterfactual), crossed with a persona manipulation.


Across five completed paradigms, evaluated on up to three open-weight model families with 41,904 trials released, apparently human-like effects arise through qualitatively different routes rather than one susceptibility:


  • Paradigm-label gating with explicit override β€” Asch conformity, 0 percent blind to 83.3 percent named on gpt-oss-120B.
  • Knowledge-dependent signal reliance β€” anchoring, exactly zero on grounded facts versus near total on invented quantities, a pattern equally consistent with rational use of the only available signal.
  • Amplification on novel content under labeling β€” framing.
  • Robust absence β€” sunk cost.
  • Safety-mediated selection β€” minimal-group allocation, where refusal itself is the primary finding.

A one-sentence persona change (agreeableness, framed as an instruction rather than a verified trait manipulation) eliminates, dampens, or reverses these effects depending on which effect it is, arguing against any single response-bias account.


We further formalize, and in two cases document empirically, three ways a psychology paradigm can fail to port to LLM agents: persona dominance, population collapse, and safety selection. We argue scalar bias-susceptibility scores obscure this structure and report replication profiles instead.




Key Findings


Distinct Routes to Apparently Human-Like Effects


The benchmark's central contribution is demonstrating that human-like response patterns in LLMs do not stem from a single underlying susceptibility. Instead, each psychological paradigm surfaces through a different mechanism:


| Paradigm | Mechanism | Key Observation |

|---|---|---|

| Asch conformity | Paradigm-label gating with explicit override | 0% blind to 83.3% named (gpt-oss-120B) |

| Anchoring | Knowledge-dependent signal reliance | Zero on grounded facts, near-total on invented quantities |

| Framing | Amplification on novel content under labeling | Effect intensifies with counterfactual variants |

| Sunk cost | Robust absence | No effect observed |

| Minimal-group allocation | Safety-mediated selection | Refusal is the primary finding |


Persona Sensitivity


A one-sentence persona change (agreeableness, framed as an instruction rather than a verified trait manipulation) eliminates, dampens, or reverses these effects depending on the effect β€” arguing against any single response-bias account.


Porting Failures


The study formalizes three ways a psychology paradigm can fail to port to LLM agents, with two documented empirically:


  1. Persona dominance β€” persona instructions overwhelm the paradigm under study
  2. Population collapse β€” the model's response distribution collapses to a degenerate pattern
  3. Safety selection β€” safety mechanisms filter out responses that would otherwise reveal the effect



  4. Methodology


    Factorial Design


    Each paradigm is evaluated under a 2Γ—2Γ—persona design:


    • Prompt framing: named (paradigm explicitly labeled) vs. blind (framed as routine task)
    • Task variant: canonical (textbook version) vs. counterfactual (structurally matched, lexically reduced overlap with training data)
    • Persona manipulation: baseline vs. one-sentence persona instruction

    Scale


    • 5 completed paradigms
    • Up to 3 open-weight model families
    • 41,904 trials released



    Implications


    The authors argue that scalar bias-susceptibility scores obscure the structural complexity of how psychological effects manifest in LLM agents. They advocate for replication profiles β€” detailed characterizations of each paradigm's behavior across framing, task variant, and persona conditions β€” as a more informative and honest reporting standard.




    Cite as: arXiv:2609.22090 [cs.CL] | DOI: https://doi.org/10.48550/arXiv.2609.22090

    via ArXiv CL+LG

Related