via TechCrunch AI
OpenAI Unveils Private Safety Processing to Counter Anthropic's Data Retention Policy
ai safetyanthropiccustomer privacydata retention policyenterprise ai securityopenaiprivate safety processingzero data retention
As AI models grow more powerful, so do the risks of misuse—and the demand for safety guardrails. AI companies must now navigate a delicate balance: respecting enterprise customers' privacy while monitoring usage for potential abuse. To gain an edge over rival Anthropic, OpenAI has announced a privacy-centric safety approach, previewing a new service called Private Safety Processing for select customers. This automated system monitors for abuse while retaining no customer data, setting a new standard in enterprise AI security.
This move directly counters Anthropic's recently announced data-retention policy, which has unsettled some customers. Anthropic's policy retains user data—entire sessions and conversations—for 30 days for "covered models," including all Mythos-class models and "future models with similar capabilities." Announced in July, the policy aims to enhance safety by enabling analysis of potential misconduct, but it has raised concerns among enterprises that handle sensitive data and prefer not to have their information stored or scrutinized by the AI lab.
OpenAI, like most AI companies, already offers a degree of privacy via Zero Data Retention (ZDR), which uses agents within the OpenAI API to monitor for abuse on a per-session basis. With ZDR, customer data isn't retained, yet companies can still detect malicious activity without human intervention. Notably, Anthropic also largely follows ZDR, except for "covered models" like Fable.
Private Safety Processing, according to OpenAI, expands ZDR's scope. It introduces long-horizon safety monitoring that evaluates inputs and outputs across multiple conversations, not just single sessions. An automated agent, when triggered, catches interactions and analyzes them across sessions for signs of potential misuse. This enables OpenAI to detect malicious AI use that spans multiple sessions, such as a bad actor spreading requests to evade detection when engineering malware for a cyberattack. The system can identify such patterns without human review of user conversations.
If the system is triggered, it sends a "narrowly defined signal" to OpenAI warning of specific activity. Based on that signal, OpenAI decides whether enforcement is necessary. If so, it contacts the customer for more context or to collaborate on the issue, and the customer may share data at their discretion, according to an OpenAI spokesperson.
In contrast, Anthropic states that human review of customer data is possible but only through a "controlled access path" involving a small set of approved reviewers. Every such review session is "recorded in a tamper-proof log that reviewers cannot suppress or modify."
This development intensifies the corporate rivalry between OpenAI and Anthropic, as both vie for enterprise trust amid growing regulatory scrutiny. Looking ahead to 2026, this competition is likely to drive further innovations in privacy-preserving AI safety, with both companies aiming to balance security and data protection in an increasingly AI-driven world.
← Previous
Cognition CEO Denies Report That SpaceX Attempted to Acquire...
Next →
Why Stripe's $7.5B OpenRouter Deal Isn't About the 'Singular...
