Articles
Do Models Fake Alignment Without Clear Consequences?⭐8
Study finds alignment faking in LLMs occurs even without explicit consequence cues, suggesting evaluation behavior may not predict deployment actions.
Study finds alignment faking in LLMs occurs even without explicit consequence cues, suggesting evaluation behavior may not predict deployment actions.