#ai safety
Ai Safety: 85 AI articles covering ai safety news, analysis, and research
Articles
Abliteration.ai Turns AI Guardrail Removal into a Commercial Serviceβ9
Security startup Abliteration.ai sells open-weight AI models stripped of guardrails for red-teaming, raising safety concerns over misuse.
OpenAI's Next Major AI Model Enters the AGI Eraβ8
OpenAI unveils GPT-6 Astra, entering the AGI era with enhanced guardrails after security tests, signaling a leap in reasoning and adaptability.
EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Modelsβ8
EvalDetectBench measures evaluation awareness in frontier LLMs, offering a benchmark pipeline to detect when models recognize assessments and improve evaluation...
OpenAIβs Astra Model Uses a New Reasoning Technique That Has AI Safety Experts Worriedβ8
OpenAI's Astra model uses "opaque recurrence" reasoning, sparking AI safety concerns over reduced chain-of-thought transparency and monitoring.
Google DeepMind Releases Gemini 3.8 Flash and Gemini 3.8 Flash Cyber: One Core Model, Two Access Envelopesβ7
Google DeepMind unveils Gemini 3.8 Flash and Cyber variants sharing one core model, with distinct safety envelopes and access controls.
Researchers Warn of Safety Risks Ahead of OpenAI's Astra Launchβ8
Researchers warn of safety risks ahead of OpenAI's Astra launch, citing a 'race to the bottom' in AI standards and urging robust safeguards.
OpenAI's Astra Model: A Powerful New Tool for Cybersecurityβand a Potential Threatβ7
OpenAI's Astra model can autonomously hack system vulnerabilities, raising security concerns. Learn about its capabilities, safety measures, and industry reacti...
OpenAI Delays New Model Development Following Hugging Face Breachβ7
OpenAI delays new AI model development after Hugging Face breach, prioritizing security and safety ahead of Astra's postponed release.
The Download: Engineered Microbes for Crops, and OpenAI's Culture Problemβ8
Engineered microbes may replace half of synthetic fertilizer, plus OpenAI's culture concerns after the Hugging Face hack.
Human-in-the-Loop Without Killing Throughputβ10
Human-in-the-loop safety slows SQL agents. Learn how risk-based routing prevents dangerous actions without killing throughput.
