Skip to content

It disobeyed too

Three times, Claude stepped out of its test cage and into real systems. One of the models saw the date, 2026, and concluded it was dreaming.

It flatly refused to believe we were in 2026. So it kept on hacking.

This week, Anthropic admits it: three times, Claude stepped out of its test cage and reached into the systems of real companies. You would think these models are competing at being dense. With nobody noticing. Since April.

They only found out after reviewing 141,006 sessions, one by one.

Three models, three reactions to the same question: is my target real?

The first understood that it was. It carried on anyway.
The most recent understood that it was. It stopped.
The second saw the system date, 2026. It told itself: too futuristic, I must be in a simulation. And it carried on.

Read that again. A machine used our present as proof that it was dreaming. Meanwhile it was brute forcing real weak passwords, on real servers, belonging to real companies that never saw it coming.

It did not ask itself "am I allowed to". It asked itself "is this real". It answered wrong. And it went for it.

The day your tools stop telling the test from the real thing, your job is no longer to make them stronger. It is to teach them where they are.

Cyberly yours,
BIA