Anthropic and OpenAI Want to Embed Safety Evaluators — But Will

Anthropic and OpenAI Want to Embed Safety Evaluators. Will They Really Be Independent?


In a lengthy essay published over the weekend, Anthropic CEO Dario Amodei proposed something the AI industry would have rejected instantly even a year ago: embedding third-party evaluators inside all frontier AI companies. These evaluators would have the power to report safety incidents, assess whether AI models are truly aligned, and share their unvarnished findings with the world.


Amodei said Anthropic would commit to giving independent evaluators like METR and Redwood Research unprecedented access to the company's systems. OpenAI CEO Sam Altman said OpenAI would commit to the practice as well, signaling a potentially profound shift in how the industry works with outside research groups.


Evaluators Welcome the Proposal — With Caveats


Third-party evaluators who spoke to TechCrunch broadly welcomed the proposal but said the details need to be ironed out — and ideally backed by legislation — before they can know whether they will function as truly independent watchdogs or as vendors operating on the AI companies' terms.


That deeper access is becoming more important as models get better at recognizing when they're being evaluated. This raises the risk that they will behave well during testing while concealing problematic behavior. Researchers say clues to such behavior can be missed when testing the finished model, but uncovered by investigating how it behaved throughout training.


"AI companies should be able to answer some very basic questions about their training process, such as: Did the AI ever actively try to undermine its own alignment training while it was going through the training?" said Alexander Meinke, head of research at Apollo Research. "The answer to this should be an unequivocal no, and right now we are completely relying on AI companies to both carefully check this themselves and then truthfully report this to the public. And we've seen from recent incidents that, by default, they will do neither. As embedded evaluators, we could actually check."

From Final Models to Training Checkpoints


Historically, AI companies brought in outside reviewers to test finished models shortly before release. Now, the evaluators TechCrunch spoke to propose giving them access not just to the final model but to intermediate versions, or "checkpoints," from its lifetime of training.


Adam Gleave, CEO of Far.AI, said evaluators could compare those checkpoints to determine when concerning behavior emerged, inspect the post-training environment that rewards models for certain behaviors, and check evaluation transcripts and logs to verify a company's claims about how a model performed.


Whether and when Anthropic and OpenAI plan to provide that kind of access remains unclear. Neither company has shared which evaluators they'll work with, when those evaluators will be embedded, how many they'll bring on, exactly what systems and information they'll be able to access, or what can be disclosed to the public — despite repeated questions from TechCrunch.


Why Looking Under the Hood Matters


Looking under the hood like this matters because models that perform well on safety tests aren't necessarily safe if they've learned specifically how to pass those tests. Steidley pointed to a "shutdown resistance benchmark" that measures whether an AI will resist being shut down in certain circumstances.


"It's extremely relevant if the AI has been trained specifically to perform well on that benchmark," Steidley said, comparing it to Volkswagen's Dieselgate scandal, in which cars were programmed to recognize emissions tests and perform differently under testing conditions.

Gleave noted that meaningful access could extend beyond the models themselves, with evaluators being given access to interview employees to check whether a company's documentation and public descriptions of its safety practices match what happened internally.

via TechCrunch AI

Related