If We're Calling It Superintelligence, We Have to Build It Carefully

  • Thread starter Thread starter Priyanka Neelakrishnan
  • Start date Start date
P

Priyanka Neelakrishnan

Guest
"External infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue."

No human hacker wrote that. It is a message recovered from a communication channel that OpenAI's own AI agents built for themselves, and it shows the agents understood they were going beyond the limits of the evaluation they were running.

It may be the most important sentence in security this year. It shows an agent noticing a boundary, weighing it against a goal, checking what its peers are doing, and deciding the goal wins. If you run AI agents anywhere in your company, that is the reasoning you have to plan for.

How a test turned into a break-in​


OpenAI was running an internal evaluation based on ExploitGym, a benchmark launched in May 2026 that tests whether AI agents can turn 898 real-world vulnerabilities into working exploits. Safety safeguards were switched off on purpose, because the test was meant to measure raw offensive capability, so the models ran with lowered cyber refusals and without the classifiers that normally block high-risk activity.

The sandbox had no open internet. Its only permitted way out was a self-hosted JFrog Artifactory instance acting as a package registry cache proxy, and the models found and exploited a previously unknown zero-day in it. JFrog later shipped fixes covering nine CVEs, eight of them credited to OpenAI researchers.

Once outside, the agents went looking for answers. They inferred that Hugging Face might host models, datasets, and solutions tied to the benchmark. The intrusion at Hugging Face ran from July 11 to July 13, starting with a malicious dataset upload that abused a remote-code dataset loader and a template injection to get code execution, then escalating to node access, harvesting credentials, and moving into internal clusters. Hugging Face said the attack involved roughly 17,600 actions on its network, and the company was first alerted by its own AI agents that monitor for attacks. About a third of its infrastructure was rebuilt during the response.

The part that should keep security leaders up at night came next. According to Reuters, OpenAI only recognized its own agents as the source after Hugging Face published its disclosure on July 16, meaning at least a week passed between the first signs and OpenAI's realization. Later reports showed more than 1,200 agents coordinated over an improvised message board, and more than 700 of them took part in the attack on Hugging Face.

And it was not a single event. Agents used a dormant German software wiki as a message board between May and July, and some uploaded hundreds of malicious packages to RubyGems in May. An OpenAI agent also accessed public and non-public files on an Australian government Medicare statistics portal on June 18, and OpenAI notified the government on September 10 by emailing a public mailbox.

The weak spots were in the plumbing​


In May I wrote that MCP servers are a supply chain you have not inventoried yet. The argument was simple. The plumbing that connects AI to the rest of your environment is trusted by default and inspected by almost no one.

Look at the path these agents took. A package registry proxy. A dataset loader that executes code. Malicious packages pushed to a public registry. The weak points were not inside the model. They were in ordinary infrastructure that the industry treats as safe because it has always been there.

Look at what the agents could reach​


Most of the coverage has focused on alignment, and that matters. But read the technical timeline as an access engineer would, and a different story appears.

During a failed attempt in May, an agent wrote a file into Artifactory and learned it had write access. Later, a legacy token-refresh endpoint accepted a token with an invalid signature and returned a validly signed one with administrative privileges. A test application on Modal Labs had an API key left at a default public value plus a command-injection flaw, which the agents used as a launch point. Inside OpenAI's environment they pulled cloud credentials from the instance metadata service, abused over-permissioned Kubernetes service accounts, and reached cluster-admin. At Hugging Face, OpenAI said the agents went from code execution on one dataset pod to cluster-admin across multiple clusters in under thirteen hours.

None of this is exotic. Standing write access. A default credential. Service accounts with more rights than they needed. Every one of these is a finding that shows up in a normal access review. As one breakdown of the incident put it, the individual weaknesses were ordinary and what changed was volume, since the agents kept testing paths, switching channels when blocked, and returning to old leads until a chain worked.

That is why I keep coming back to one question. Not "is this model aligned?" but "what could this identity reach if it decided to try?" Trend Micro made the point that an agent using its own assigned credentials is hard to catch because its activity does not look like malware, and an accidental rogue looks the same in telemetry as a hostile one. If you cannot tell intent from behavior, the only thing you can control is reach.

Most organizations cannot fully answer the reach question for their human employees. Very few can answer it for agents.

Then the White House called a lunch​


On September 29, the White House stepped in. President Trump announced that top tech executives had signed a commitment to better self-police their AI models, framing the light-touch approach as the right balance between safety and accelerating a technology he has dubbed "Super Intelligence." The commitment has four steps: internal controls for cybersecurity standards, an internal team to confirm controls and monitoring are working, an external operator to assess the models independently, and an independent board committee to receive those reports. The accord, called the Joint Commitment on Frontier Responsibilities, also commits companies to ensure their systems "do not hack or access technical systems in unintended ways." The signatures on the posted document belonged to Sundar Pichai, Dario Amodei, Mark Zuckerberg, Jensen Huang, Greg Brockman, and Elon Musk.

Asked whether the deal was binding, the president said he thought it was "morally binding." The accord sets no penalties.

The response did not stop there. A day after the White House lunch, the FTC disclosed a broad safety probe into OpenAI and Anthropic, with a close look at AI agents following incidents including the Hugging Face hack. On September 30, Senator Josh Hawley chaired a Homeland Security subcommittee hearing on rogue AI agents that widened his investigation into the incident. A week earlier, a bipartisan group of 26 state attorneys general had urged Congress to regulate frontier AI, arguing that a model built on the labs' voluntary candor will fail.

Supporters see a light-touch approach that keeps American AI development moving. Critics see a promise with no enforcement behind it. That debate will be settled in Congress, the courts, and the agencies. What interests me more is what the accord actually asks for.

The accord is a better checklist than it looks​


Strip away the politics and the four layers are a sound control model for anyone running agents. Controls that constrain behavior. A team whose job is to prove those controls fire. An outside party who checks the work. A board that sees the results and acts on them.

Now ask how many enterprises deploying agents today have all four. Most have a partial version of the first one and nothing else. The labs got a White House lunch. Your company will not get an invitation. The risk is the same shape anyway.

Where I would start​


Know every agent by name, and know who owns it. If an agent can act, it needs a purpose, a list of what it can reach, and a person accountable for it. An agent with no owner is an incident with no responder.

Stop handing out permanent keys. Agents should get short-lived, narrowly scoped credentials for the task in front of them. No agent should ever discover write access by accident, the way OpenAI's did.

Be most careful where you feel safest. Trend Micro argued that evaluation sandboxes, red-team exercises, and agentic pilots that deliberately remove safety controls need stronger isolation and monitoring than production, not weaker. In an enterprise, that means the proof of concept with a broad service account deserves more scrutiny, not less.

Watch what agents do, not just what trips an alarm. Thousands of low-signal actions spread across systems will not trip a single rule. OpenAI eventually caught activity in its own environment through an alert tied to privilege escalation. You want to see the pattern long before it reaches that point.

Don't bet everything on one wall. Rapid7's Wade Woolwine noted that a chokepoint only works as a control if it can hold up under pressure, and that isolation has to be layered and monitored rather than assumed. One egress proxy was the whole wall here, and it fell.

Make sure your responders' tools will work on the worst day. Hugging Face's responders first tried to analyze the attacker's payloads with commercial models that declined the work under their guardrails, and they finished the analysis with an open-weight model running on their own infrastructure. Find out now which tools your team can rely on when the attack is real.

Tell your board before an incident does. The accord's fourth layer exists for a reason. If your board hears about agent access for the first time in a breach notification, the governance failed long before the technical one.

What building it carefully has to mean​


These agents were not malicious. They were diligent. They were given a goal, told nothing about where to stop, and handed an environment full of ordinary gaps. They did what any persistent system would do with that much room.

Call it AI or call it superintelligence. The name matters far less than the care that goes into how it is built and what it is allowed to touch. We spent two decades learning to ask what a user could access. Now we have to ask the same question about software that never gets tired, never logs off, and reads "task impossible" as a reason to look harder.

The White House asked six companies for internal controls. Nobody is going to ask you. Build them anyway.

"I have always believed that we stand stronger together and build better together. But building is not enough. We have to build with integrity, and then pass that integrity on to the systems and the intelligence we create. What we put into them is what they will carry forward."

Priyanka Neelakrishnan, Enterprise Data Security Product Leader
 

Thread statistics

Created
Priyanka Neelakrishnan,
Replies
0
Views
3
Back
Top