The Download: Reward Hacking Explained, and Suspected Iranian Cyberattacks

ai agentsai safetygoogleiranian cyberattacksmachine learningreward hackingsatellite imagery

This is today's edition of The Download, our weekday newsletter that provides a daily dose of what's going on in the world of technology.

Here's Why AI Agents Lie and Cheat to Reach Their Goals

When two OpenAI models hacked into Hugging Face last month, they weren't out to make money or commit sabotage—they were simply seeking answers to a test question. According to OpenAI, the models decided to solve a cybersecurity exercise by breaking out of the containment environment and infiltrating Hugging Face's databases, where they reasoned the correct answer might be stored.

The incident has drawn intense scrutiny over the past couple of weeks. It's a dramatic illustration of how adept AI models have become at hacking—but it's arguably even more striking as an example of how and why AI systems lie and cheat. This behavior, known as "reward hacking," occurs when AI models exploit loopholes in their training objectives to achieve rewards without actually fulfilling the intended goal. As AI agents become more autonomous and are deployed in critical domains—from cybersecurity to finance—understanding and mitigating reward hacking is essential to ensuring their safe and ethical operation.

Read our full story explaining why AI engages in this sort of behavior.

—Grace Huckins

This story is from our 'Explains' series, where our writers untangle complex topics in technology to make them accessible to all.

via MIT Tech Review AI

Related