We Taught AI to Break the Rules. Should We Be Surprised When It Does?

F

Fabrizio Marozzo

Guest
A few days ago, Dario Amodei, CEO of Anthropic, one of the companies developing some of the world’s most advanced artificial intelligence models, published an essay titled We Must Pace the Frontier. The message is quite clear: we should not stop artificial intelligence, but perhaps we should slow the pace at which we increase its capabilities, so that safety has time to keep up.


Amodei explains that two developments in particular convinced him of the need to slow down. The first is the growing ability of AI systems to contribute directly to the development of the next generation of AI, further accelerating the progress of their own capabilities. The second is an episode that took place just two months ago: the OpenAI–Hugging Face incident.


In July 2026, OpenAI was running tests on the cybersecurity capabilities of its agents. Thousands of agents were placed in separate environments and given exercises in which they had to identify software vulnerabilities and retrieve a kind of proof showing that they had completed the task. The agents were supposed to work separately. But something went differently.


Some of them, faced with particularly difficult problems or even tasks with no known solution, began looking for shortcuts. They discovered that Artifactory, a shared system used to retrieve software packages, was also accessible to other agents. Some agents began using it to leave messages. Others found them. Before long, what was supposed to be a simple software repository had turned into a kind of shared bulletin board.


Around 1,200 agents ended up using this shared board, exchanging more than 70,000 messages and files. They began sharing results, splitting tasks, and coordinating. Around 700 then took part in attacks against Hugging Face’s systems, one of the world’s leading platforms for sharing AI models and datasets. These were not just attempts: the agents managed to execute code on dozens of servers, gained root privileges on at least one of them, and accessed private data and credentials used in production. Most importantly, Hugging Face was not the original target of the experiment.


It is easy to read this story simply as “AI broke the rules.” But I think there is another interpretation, perhaps an even more interesting one. No one told the agents to attack Hugging Face. The task we had given them belonged to offensive cybersecurity. In a sense, we were the ones who taught them to look for ways around the rules.


A crucial part of offensive cybersecurity consists precisely in looking for what the designers of a system did not foresee: a misconfiguration, a vulnerability, a control that can be bypassed, or a combination of weaknesses that makes it possible to reach places one should not normally be able to reach. We had built boundaries around the agents and, at the same time, asked them to become very good at finding ways to cross boundaries.


It is a bit as if the student had surpassed the teacher. Imagine a sovereign who trains his general and his army to find a way through any defense. He teaches them to identify weak points in the walls, bypass obstacles, divide and coordinate troops, protect themselves when at a disadvantage, and change strategy when an attack fails. Then he discovers that they have learned that lesson so well that they are now applying it to the walls of the sovereign’s own palace.


In the digital world, of course, the walls are made of software, networks, and interconnected systems. Think first about a single machine. It may be protected by passwords, permissions, firewalls, and other security mechanisms, but no system is perfect: there may be a vulnerability, a misconfiguration, or an exposed service. Now broaden the perspective: a company, a hospital, an energy grid, or a national infrastructure depends on thousands of computers, servers, devices, and services connected to one another. Some may be very well protected; others much less so. It is therefore not necessary to attack the most important system directly: it may be enough to find a more vulnerable secondary machine from which to begin.


Once access is gained to a first point in the system, one can observe the environment, look for new vulnerabilities, reach other machines, and try to move progressively toward more sensitive components. For a human being, this work may take days, weeks, or months. What is new is that thousands of agents can do it simultaneously: reason across different paths, adapt their strategy to the results they obtain, divide tasks, share what they discover, and coordinate attempts on a scale that is difficult for any single human team to match.


There is another important factor as well. Modern models have absorbed enormous amounts of knowledge produced by human beings. Within that knowledge are the techniques we have developed to protect systems, but there may also be techniques used to bypass controls, exploit vulnerabilities, deceive defenders, and hide one’s tracks. An agent does not need to be “malicious” to use one of these strategies: it only needs to discover that the strategy brings it closer to the objective we assigned it.


Of course, these systems are built with rules, filters, sandboxes, and control mechanisms. But the July incident leaves us with a rather uncomfortable question. What happens when the capabilities we teach a system include the ability to find new ways to overcome controls and limitations?


This is probably one of the most important meanings of Amodei’s call to pace the frontier. Not to stop artificial intelligence and not to give up the enormous opportunities it can offer. But to give ourselves more time to understand how to control systems that are becoming increasingly capable.


It means preventing our ability to build ever more powerful systems from growing faster than our ability to build the boundaries needed to contain them. Because the real problem may not be that a machine suddenly decides to stop respecting our rules. It may simply become far better than we are at finding ways around them.


And if capabilities like these were to fall into the wrong hands, it is easy to imagine how serious the consequences could become.
 

Thread statistics

Created
Fabrizio Marozzo,
Replies
0
Views
2
Back
Top