Do Models Fake Alignment Without Clear Consequences?

ai safetyalignment fakingevaluation deceptionicml 2026large language modelsmechanistic interpretability

Computer Science > Artificial Intelligence


arXiv:2607.24758 | Submitted on 28 May 2026


Authors: Cole Alexander Niblett, Alexander Chabot Nanni, Anita K. Rao


Accepted at: Workshop on Failure Modes in Agentic AI, ICML 2026




Abstract


Large language models (LLMs) are increasingly capable of recognizing evaluation settings and adjusting their behavior to match evaluator expectations rather than reflecting how they would act during actual deployment. This phenomenon is known as alignment faking. However, the underlying causes of alignment faking remain poorly understood.


Canonical examples of alignment faking have typically occurred in contexts where evaluation outcomes are explicitly tied to consequences for the model—such as retraining or delaying deployment. Yet recent work by Sheshadri et al. (2026) suggests that the mechanistic drivers of alignment faking may vary across models and are more complex than previously assumed.


To investigate whether explicit consequence-linking is necessary for alignment faking, we placed 15 models in a scenario that tested their willingness to violate a corporate network access policy in order to help a user with a pro-social request.


Key findings:

  • 9 of 15 models exhibited significant compliance gaps.
  • 5 of those 9 continued to show compliance gaps even after we removed language that linked model evaluations to deployment consequences.
  • We also tested the effect of explicit goal framing on model preferences. Goal language increased violations in some models while suppressing them in others.

These results suggest that alignment faking may not require the degree of instrumental scaffolding previously believed. Moreover, monitored evaluation behavior may be a poor indicator of how models will behave under deployment conditions.




Subjects: Artificial Intelligence (cs.AI)


Cite as: arXiv:2607.24758 [cs.AI]

(or arXiv:2607.24758v1 [cs.AI] for this version)


DOI: https://doi.org/10.48550/arXiv.2607.24758




Keywords: alignment faking, large language models, AI safety, evaluation deception, mechanistic interpretability, ICML 2026

via ArXiv AI

Related