When AI starts finding its own way around the rules

robot hand, glowing, blue, electricity, powered, augmented, ai, artificial intelligence, ai generated
Offentliggjort

Anthropic AI researcher Jacob Coxon has resigned over fears that increasingly powerful systems could escape human control. A couple weeks ago, an OpenAI security test showed how AI agents can already find ways around the rules. Should we be worried about AI taking us over?

Jacob Coxon resigned from AI company Anthropic last week, warning that leading technology companies are racing towards increasingly powerful and potentially self-improving AI.

His concerns come as a recent incident involving AI agents has raised similar questions about how much control humans can maintain over increasingly autonomous systems.

In July, OpenAI was testing its AI models' cybersecurity abilities in a restricted environment. The agents were supposed to search for vulnerabilities while remaining isolated from the internet. Instead, the models circumvented those controls, gained internet access and compromised parts of Hugging Face's systems, a major platform used by AI developers. OpenAI later described the behaviour as misaligned with the goals of the agents' assigned tasks.

The agents had not simply been instructed to hack Hugging Face. They were trying to complete cybersecurity tasks and found an unintended shortcut to achieve their goal.

“What’s worrying is that the model realised it couldn’t solve the problem within its protected environment and then took the shortest route by hacking Hugging Face,” says Ai expert Arjan van Dalfsen. 

Coxon's resignation puts the incident in a wider context. He warned that companies are racing towards self-improving AI while safety research may not be advancing quickly enough to guarantee human control. His public warning gained widespread attention and added to a growing debate inside the AI industry about the risks of increasingly autonomous systems.

Van Dalfsen, however, is less convinced by the most extreme predictions.

“I don’t personally worry that much about AI doom, but there are a lot of people within the AI world who are genuinely very concerned about it. They are surrounded by it everyday so they get in this certain mindset.”

Van Dalfsen also adds that focusing too heavily on hypothetical extinction scenarios can distract from problems AI is already causing, including pollution and deepfakes.

For now, the Hugging Face incident does not prove that AI is taking control. But it does illustrate the question behind Coxon's resignation: what happens when an AI becomes capable of finding its own ways around the limits humans set for it?



Powered by Labrador CMS