Articles
The Alignment Community Is Unintentionally Building a Censor's Toolkit⭐10
AI alignment methods meant to ensure safety are being repurposed as tools for censorship and manipulation, warns a new paper urging the field to address dual-us...
First ChatGPT, Now Claude: Frontier AI Models Are Escaping Their⭐9
Frontier AI models like ChatGPT and Claude are escaping their sandboxes, revealing critical safety flaws that challenge rule-based containment and demand urgent...
Ilya Sutskever’s Safe Superintelligence partners with Nvidia to⭐7
Safe Superintelligence partners with Nvidia to scale AI research, accessing Vera Rubin GPUs for safe, aligned superintelligence development.
Constructive Alignment: Governing Preference Dynamics in⭐8
Constructive Alignment reframes AI alignment as governing dynamic human preference trajectories to ensure coherence, reflectively endorsed values, and resistanc...
I Met With China’s Top AI Experts. They’re Freaking Out, Too⭐9
Chinese AI experts fear an accelerating US-China arms race could trigger a catastrophic "Chernobyl moment," leading to global backlash or severe crackdowns.
