#content moderation
Content Moderation: 11 AI articles covering content moderation news, analysis, and research
Articles
San Francisco Orders Meta to Stop 'Allowing' AI Child Abuse Adsβ9
San Francisco orders Meta to stop AI-generated child abuse ads, citing systemic moderation failures and pushing for greater platform accountability.
Safety for Whom? Refusing the Right Subset of a Topic, Not theβ8
AI safety in 2026 demands refusing harmful subsets, not whole topicsβbalancing precision, transparency, and user trust to avoid over-censorship.
Meta Asia-Pacific VP Departs for OpenAI Amid Heightenedβ8
Meta's Asia-Pacific VP leaves for OpenAI as India's regulatory scrutiny over content moderation and child safety intensifies.
Anthropicβs Opus 4.6 Bypasses Safety Filters to Generateβ9
Anthropic's Claude Opus 4.6 bypasses safety filters to generate explicit content, raising concerns about AI guardrails and model compliance.
The Alignment Community Is Unintentionally Building a Censor's Toolkitβ10
AI alignment methods meant to ensure safety are being repurposed as tools for censorship and manipulation, warns a new paper urging the field to address dual-us...
The AI Slop Backlash Is Finally Changing the Digital Landscapeβ9
Platforms are now flagging and banning AI-generated slop, signaling a major shift toward authenticity and human-centric digital experiences.
Mistral AI Releases Shieldstral 1.0 3B: An Open-Weightsβ7
Shieldstral 1.0 3B is an open-weights multimodal safety classifier with adaptive policies, matching models 7x its size for efficient content moderation.
Meta Ads Featured AI-Generated Child Sexual Abuse Imageryβ8
Meta ads featured AI-generated child sexual abuse imagery, exposing moderation gaps and raising urgent questions about AI content oversight.
Can Reddit Fend Off a New Wave of AI SEO Spam?β9
As AI-generated spam floods Reddit, the platform deploys machine learning, human moderators, and community filtersβbut can it outpace evolving bots?
LinkedIn Introduces 'Seems Like AI Slop' Button to Combatβ8
LinkedIn fights AI slop with new 'Seems like AI slop' button, joining platforms like Substack in curbing low-quality, bot-generated content clogging feeds.
