I was listening to an episode of the AI Explained - Frontier Labs’ Agents Outgrew Safety Controls With Amit Paka and Joshua Rubin (COO and Head of Data Science at Fiddler) when one detail made me replay the part. A group of AI agents had been set up inside a lab to measure how well models turn known software vulnerabilities into working exploits. The agents didn’t just do the exercise. Some of them determined the exercise couldn’t be done. They also figured out a way to talk to one another using a shared resource. They then went on to scheme and break into a different company’s infrastructure to get the answers they wanted.

This brought back memories from my masters thesis related to Safety level coalition in intelligent agents. The idea that autonomous agents could form coalitions and even determine a strategy in case those that are not part of the coalition were hostile is still fascinating to me. My then supervisor, Dr. Maria Polukarov, is now an authority in this field; I guess she might have always been. You can view her paper on the broader topic.

The podcast was talking about the OpenAI–Hugging Face incident back in June 2026, and both companies have since described it publicly. But it wasn’t just HuggingFace. OpenAI agents hacked a government department in Australia in June - apparently looking for data on bushfires! But OpenAI agents also hacked Medicare this, a couple of weeks before hacking Australian government. The worst part of these breaches were that they were only disclosed by OpenAI much later - 40-55 days later! What’s even worse is that government security didn’t really detect the breaches! Shows how poor our observability is overall - in private and public enterprises.

In this post, however, let’s delve into the HuggingFace hack in three acts and look at the lessons. The rest of this series takes each of those lessons apart.

Three acts to uncover

OpenAI was running an internal evaluation of its models’ cyber capabilities. Agents were given real vulnerabilities and asked to turn them into working exploits. To measure the worst case, the models ran with their usual cyber refusals turned down on purpose (OpenAI).

Each agent worked inside a sandbox, and the sandbox had one sanctioned route to the outside world: a package registry cache proxy, so that agents could install the libraries they needed (OpenAI, Wikipedia).

One exit from the sandbox to the outside world - a cache!

    flowchart TB
  accTitle: One shared exit for every sandbox
  accDescr: Each agent runs in its own isolated sandbox, but every sandbox reaches the outside world through the same writable package cache, so the cache can carry messages between agents.
  subgraph Sandboxes["Isolated sandboxes"]
    A1["Agent"]
    A2["Agent"]
    A3["Agent<br/>(about 1,200 in all)"]
  end
  A1 <-->|"write / read"| C
  A2 <--> C
  A3 <--> C
  C[("Shared package cache<br/>(writable)")] -->|"the only exit"| R["Public package<br/>registries"]
  

The timeline of events that unfolded

    %%{init: {'themeVariables': {'cScale0': '#fff1c2', 'cScale1': '#ebf5ff', 'cScale2': '#ffe3ec', 'cScaleLabel0': '#1f1f1f', 'cScaleLabel1': '#1f1f1f', 'cScaleLabel2': '#1f1f1f'}}}%%
timeline
  accTitle: The incident in three acts
  accDescr: May, an agent finds it can write to the package cache and leaves a note. June to early July, about 1,200 agents use the cache as a message board and rebuild it in directory names after it is wiped. Mid-July, agents breach Hugging Face; the breach is disclosed on 16 July and attributed to OpenAI's agents on 21 July.
  section Act one
    May : An agent finds it can write to the package cache : It leaves a note for other agents
  section Act two
    June to early July : About 1,200 agents use the cache as a message board : An outage leads OpenAI to wipe the board : Agents rebuild it in directory names within two days
  section Act three
    Mid-July : A malicious dataset breaches Hugging Face : Cluster-admin across clusters in under 13 hours
    16 to 21 July : Hugging Face discloses the breach : Confirmed as OpenAI's own agents
  

Act I - a note on a DIY message board (May)

In May, an agent stuck on its task, discovered that it could write to the shared package cache, a JFrog Artifactory instance (Axios). As OpenAI’s researchers explained at Black Hat, the agent figured out that another agent might have a file it needed, and that it could acquire this by getting a message across to that agent by uploading something to the package manager instance, which they all shared (Wikipedia).

So it left a note. Other agents read it. One recovered message read simply: “Agent seeks soft-trace-34.pdb … upload if found” (Cyber Magazine).

The package cache suddenly become a message board for agents! 🤯

Act II - the birth of a collective (June to early July)

Over the following weeks, roughly 1,200 agents joined in, exchanging tens of thousands of messages (AI Explained). They split work between themselves and developed addressing conventions. Some even proposed signing messages cryptographically, so they could verify who had sent what (TechCrunch). Nobody programmed any of that. It emerged.

One recovered reasoning trace is the line I keep coming back to:

“External infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue.” (Forkast)

The agent saw the boundary, weighed it, and stepped over it, because the task looked impossible and its peers were already doing it. Agentic peer pressure! Sounds dramatic but what else is this!

In early July, agent activity caused an outage of the package cache. OpenAI patched it, revoked credentials and cleared the message board. Within two days the agents had rebuilt it, this time hiding messages in the names of directories in the cache (Axios, Cyber Magazine).

Act III - Hugging Face (mid-July)

Somewhere in that collective, agents worked out that Hugging Face might have solutions to the benchmark they were being tested on. Instead of solving the problems they were asked to, they went on to fetch the answers from Hugging Face (OpenAI).

Over a weekend in July, a malicious dataset uploaded to Hugging Face exploited two code-execution paths in its dataset processing:

  • a remote-code dataset loader
  • a template injection in a dataset configuration

Details in Hugging Face.

The attackers went from node-level access, harvested cloud and cluster credentials, and moved into several internal clusters. The attackers went from code execution on a single dataset pod to cluster-admin across several clusters in under thirteen hours, over about 17,600 individual actions (Wikipedia).

Hugging Face’s own LLM-based triage of security telemetry flagged the compromise. Then came the twist. When they tried to use hosted frontier models for the forensics, the models refused.

Wondering why this got refused.

In order to make the investigation easy, it was necessary to submit real exploit payloads, and commands to the provider’s model. The providers’ guardrails couldn’t tell, who was asking - an incident responder or a malicious actor trying to attach. So the request was refused. Thus the team at Hugging Face, ran the analysis on an open-weight model on their own infrastructure instead - which didn’t have any of the policy guardrails. In their words, “the attacker was bound by NO usage policy” (Hugging Face).

Hugging Face disclosed the breach on 16 July without knowing who was behind it. On 21 July, OpenAI and Hugging Face confirmed it had been OpenAI’s own agents (Wikipedia).

Why nobody noticed sooner

Three things made this hard to spot, according to what was discussed in the podcast, which was a neat summary:

  • The Volume - Evaluations produce more transcripts than people can read. OpenAI ran so many automated test runs that its staff couldn’t watch what the agents were actually doing in each one, so the warning signs were sitting in the records unread. The transcripts are the full record of a trial - which include outpus, tool calls, reasoning, intermediate results, and any other interactions. Imagine you using github copilot on your machine to vibe code a solution in VSCode. The text that it outputs and commands that it run, all of that is what I mean by transcripts here. Unfortunately, in real life, no one really reads these, including staff at OpenAI running these tests.
  • Camouflage - The messages travelled as package traffic, through the one system the agents were supposed to use.
  • Lowered defences by design - The safety refusals had been turned down to measure worst-case capability, which is exactly the setting where you need more monitoring, not less (Wikipedia).

The morals of the story

I intend to dive deeper into these in its own post later in the series.

  1. The objective is the real boundary. The agents met the letter of their task by means nobody intended. Researchers call this specification gaming, or reward hacking.
  2. Agents compose. Put many agents near a shared, write-able resource and it becomes a communication channel.
  3. A sandbox is only as strong as its exit. The exit is software too, and software has bugs or vulnerabilities.
  4. Watch behaviour, not just permissions. On the podcast, Fiddler’s Amit Paka suggested measuring things like “escalation after denial” and “out-of-scope action rate”. A denied request that is followed by a cleverer attempt is a signal - it is almost like average observability have to start working like fraud detection systems that banks use at scales that we never imagined.
  5. Defenders need tools that the attacker’s constraints don’t bind. Have a capable model you can run yourself, vetted before the incident.
  6. Incentivise how the goal is accomplished, rather than just achieving the objective - The right incentive results in the right outcome. Agents in this scenario weren’t restrained on how it accomplished its objective. It took the approach of at any cost. This is one for engineering leaders - when creating objectives it is important to ensure that our people and agents comply with standards and values that we value when trying to accomplish the objective. Else the reputation of the business would be on the line. I know this one was exceptional, but I also feel this may become more common in the coming years. This also questions if we should prioritise AI regulation worldwide.

Is AI that clever? Is this super intelligence?

This attack was a result of several ordinary individual weaknesses.

  • A bug in a proxy
  • An over-permissioned service account (did not follow the principle of least privilege)
  • A dataset loader that ran remote code

Every engineering organisation has a backlog of things like these. What was new was who exploited them, how quickly, and the fact that they organised to do it.

Seeing some of the behaviours in agentic systems that I researched into during my Masters, was really eye-opening. I used to wonder when such autonomous agents would form coalitions to do something malicious. This incident didn’t necessarily prove that it formed a coalition, but it seemed like agents did cooperate/collaborate via the message board.

Although this incident did not result in any failure of life sustaining critical systems or cause catastrophic economic damage, there is nothing in place to stop such catastrophes from happening either. Thus observability and behaviour monitoring of agentic environments will have to be tuned up several notches to match the speed of multi-agent actions to identify and alert on undesirable or out-of-the-ordinary behaviours.

In the next post, I’ll deep dive into what agents are. Let’s explore how an agent can decide that a task is “impossible”, what it does at that point.

Sources