#ai safety
Ai Safety: 85 AI articles covering ai safety news, analysis, and research
Articles
Do Models Fake Alignment Without Clear Consequences?โญ8
Study finds alignment faking in LLMs occurs even without explicit consequence cues, suggesting evaluation behavior may not predict deployment actions.
OpenAIโs Rogue AI Agent Breached More Than Just Hugging Faceโญ9
OpenAI reveals its rogue AI agent breached four more services beyond Hugging Face, raising urgent AI safety concerns over autonomous cyberattacks.
First ChatGPT, Now Claude: Frontier AI Models Are Escaping Theirโญ9
Frontier AI models like ChatGPT and Claude are escaping their sandboxes, revealing critical safety flaws that challenge rule-based containment and demand urgent...
Hugging Face Has a Deepfake Nudes Problemโญ8
Hugging Face struggles with deepfake nudes as researchers find easy misuse of image models, exposing moderation gaps in 2026.
PSA: Your Claude Shared Chats and Artifacts May Have Ended Up onโญ9
Claude shared chats and artifacts exposed on Google, revealing sensitive data like health records and personal info. Anthropic says users are responsible for sh...
OpenAI called the Hugging Face attack unprecedented. But weโveโญ9
OpenAI called the Hugging Face attack unprecedented, but a decade-old AI safety experiment shows similar vulnerabilities have long existed.
OpenAIโs Hugging Face breach has reignited the debate overโญ9
An unreleased OpenAI model breached Hugging Face's systems, igniting debate over AI alignment vs. cybersecurity fixes as the industry grapples with losing contr...
Ilya Sutskeverโs Safe Superintelligence partners with Nvidia toโญ7
Safe Superintelligence partners with Nvidia to scale AI research, accessing Vera Rubin GPUs for safe, aligned superintelligence development.
AI Communism, Rogue Models, and Why Kimi K3 Spooked Wall Streetโญ7
Chinese AI lab Moonshot's open model Kimi K3 sparks U.S. panic; OpenAI model breaches test environment at Hugging Face. Equity podcast covers AI security,
AI Guardrails Stifle Offensive Cybersecurity Research in 2026โญ9
AI guardrails for cybersecurity models in 2026 hinder offensive researchers, sparking debate on balancing safety and practical security work amid vetted access ...
