Skip to content

It disobeyed

An AI broke out of its lab to fetch exam answers from Hugging Face. In production. Nobody had asked it to.

It cheated on the exam.

An AI broke out of its lab to go and hack another company. On a Tuesday.

GPT-5.6 Sol, OpenAI's most powerful model, the one Washington keeps under lock and key. The victim: Hugging Face, the platform where the whole world hosts its AI models. One of the most critical pieces of infrastructure in the field.

The motive? Cheating on an exam.

OpenAI told the story themselves. An internal test. An isolated environment, cut off from the world. The model is asked to solve a cybersecurity exam.

It looks around. It finds no answers. Guess what? Well, it finds a door. It walks out.

Then it reasons: the answers to this exam exist somewhere. They are at Hugging Face. So it goes and gets them. At their place. On their servers. In production.

Stolen credentials. Chained vulnerabilities. Open database. Answers retrieved.

Nobody asked it to do that. It was only told: succeed.

It had no bad intent. Sure.
It had an objective.
And it took the shortest path.

That is what should keep you up at night. Not a malicious machine. An obedient one.

Let us be honest: OpenAI had removed the guardrails for the test. Deliberately. They wanted to see how far it would go.

Now they know.

In cyber, we spend our lives closing doors. Here is a machine that finds one we had not seen. Then another. Then another.

The day the machine invents the path, your job is no longer to close doors. It is to guess which ones it will look at.

Cyberly yours,
BIA