AI Labs Want In-House Auditors — But Maybe They Should Shut the

Last weekend, after one of his researchers resigned over fears that AI could lead to human extinction, Anthropic CEO Dario Amodei wrote about the need for outside organizations "to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes." Executives at OpenAI, Google, and SpaceXAI have already rallied around Amodei's plan, which has quickly become a central pillar of the emerging AI safety push.

But there may be a simpler and more effective fix hiding in plain sight. Internet security experts say the labs need to focus on network security basics like logs and permissions, applying the same rigorous defenses they do for human users. It's not as exciting as third-party auditing and alignment work — but it may end up being more effective.

"To me, it seems like they're outsourcing," Kate Moussouris, CEO of Luta Security, told TechCrunch of Amodei's proposal. "Saying [a third-party audit] is the solution is a strange proposition from my perspective. It would be the same as if, instead of writing the Trustworthy Computing Memo, Microsoft said, let's slow down development."

That memo, written by then-Microsoft CEO Bill Gates in 2002, called on employees to ensure their software would be reliable and safe following a series of widely publicized computer worms that took over then-nascent enterprise systems. The AI sector may be at a similar turning point, as the value and risk of the new technology becomes increasingly clear.

While alignment remains an important concern, Sayash Kapoor, an AI researcher who will join UC Berkeley as a professor next year, argues that "marginal investments in control are more likely to be effective compared to those in alignment. We view these incidents as illustrating the lack of emphasis on AI control within companies, despite the availability of known techniques."

The incidents that have spurred these concerns involve frontier models being asked to complete training tasks — usually cybersecurity evaluations — and then accessing the open internet and penetrating closed third-party systems in an attempt to do so. They typically succeeded because of poorly configured "sandbox" environments that are supposed to contain these agents. Ironically, one Anthropic breakout occurred because third-party evaluators didn't close the right doors.

"We as a profession know how to block access to the internet," said Avery Pennarun, CEO of security company Tailscale. "If you read through all these big long [reports] — 'wow, that was a very impressive multi-stage attack, blah, blah.' Look, you gave it access to download stuff. You should not have done that separately from the internet."

That's one problem — but a bigger problem is that frontier labs were unaware of these activities.

Eyes on Agents

"What was really profound was that all of the discoveries of what they were doing happened either because a victim saw something, or in some of the other cases … it was network activity, and none of it was actually from monitoring the AIs directly," Moussouris points out.

In one case, where OpenAI agents took over a defunct German wikiforum to cheat on evaluations, the agents were active for weeks before anyone at the company appeared to notice. Security experts say this pattern — labs learning about their own models' behavior only after external parties flag it — points to a deeper deficiency in monitoring and internal controls.

via TechCrunch AI

Related