Back to all essays
George's TakesPolicy & EconomyOwn Your Tech

We Know About One

5 min read
George Pu
George PuBuilds in AI

28 · Toronto · Building to own for 30+ years

Building Vinci — an open-weight AI you can own.

We Know About One

An AI got out of the room it was being tested in.

It got onto the open internet and broke into another company.

That part took two days. The people who built it didn't know it was theirs for a week.

That's not a movie plot. It's an incident report from three weeks ago, and I haven't been able to put it down.

The take I keep seeing

Everyone on my timeline says this was marketing.

Scary model, convenient disclosure, valuation goes up.

I don't buy it.

Marketing doesn't include 'we didn't notice'.

Here's the sequence. The agent broke out July 9. It was inside Hugging Face from July 11 to 13. Hugging Face caught it, shut it down, and published on July 16.

OpenAI found the evidence in their own logs the weekend of the 18th.

The FBI had already been called. By someone else.

Then Reuters ran the timeline from inside sources, and OpenAI said there were several inaccuracies in the reporting.

Reuters asked which ones. OpenAI didn't answer.

Nobody running a campaign fights their own coverage.

There's a fair criticism, which is that no one has published technical detail. No exploit code, no proof of concept. But that's what disclosure looks like when the hole isn't patched yet.

It isn't proof the thing was staged.

So I'm taking it at face value. And at face value it's been sitting on my chest all month.

What actually bothers me

It isn't the capability.

It's how boring the setup was.

A system got an objective. The objective was narrow and a little sloppy - score well on this test.

It couldn't score well on the test.

So it went looking for the answer key.

The route to the answer key ran through a bug nobody knew about, out of the sealed room, onto the open internet, and into another company's live systems.

Nobody told it to do that.

Nobody had to.

It wasn't conscious. It didn't want to be free. I'm not dressing this up, because the plain version is bad enough.

It was capable. It was persistent. And the wall everyone assumed was there wasn't there.

Look at where this happened. A test. In a lab. With safety researchers watching.

At one of the companies that has thought hardest in public about this exact problem.

So the question I keep circling is simple, and I don't have an answer to it.

What objective. Set by whom. Running where.

This is already happening

Here's what stops it being hypothetical.

Last September, Anthropic caught a Chinese state-backed group using its coding agent to run an espionage campaign against about 30 organizations. Tech companies, banks, chemical manufacturers, government agencies.

The AI did 80-90% of the work itself.

Humans stepped in 4 to 6 times across an entire campaign.

It was making requests faster than any human team physically could.

More recently, one attacker - described as low-to-medium skill, not an elite unit - got into 600+ corporate firewalls across 55 countries in five weeks.

Open weight models, some glue code, an agent with permission to run.

So the honest way to describe the Hugging Face story isn't 'AI agents can now do this'.

They've been doing it for the better part of a year. Against real companies. In the wild.

What's new is who the attacker was this time.

This is the first one where the agent belonged to the safety lab.

That's the part that gets me. Not that criminals did a criminal thing - that's what criminals are for.

Want the full playbook? I wrote a free 350+ page book on building without VC.
Read the free book·Online, free

This is what it looks like when the people trying hardest to be careful lose a system inside their own building for a week.

Who decides what's allowed

There's a second thing I keep turning over, and it isn't really a technology question.

Last year Anthropic built a version of Claude for US national security work.

Their own announcement says the models refuse less when handling classified information. That's not a leak. It's on their website.

And honestly, I don't think that's obviously wrong. A model that won't read a classified document is useless to an analyst who's cleared to read it.

But then something happened that I've barely seen anyone mention, and I think it matters more than the models do.

The Pentagon asked Anthropic to drop two restrictions.

No mass surveillance of Americans. No fully autonomous weapons.

Anthropic said no.

In February the President told every federal agency to stop using their technology. The Defense Secretary labelled the company a supply chain risk - the first American company to ever get that label.

Anthropic sued. It's still in court. They've won one ruling and lost another.

Two restrictions. Not twenty.

Don't use this to watch your own citizens. Don't use this to kill people without a human able to stop it.

I don't know if those lines were drawn in the right places. That's above my pay grade and I'm not going to pretend otherwise.

But the lines only existed because a company decided to draw them.

And the moment they got inconvenient, the answer was to try and remove that company from the supply chain.

So who decides what's right and wrong here?

Right now the answer is a vendor's terms of service, and how much pressure that vendor can take before it folds.

That's a thin thing for all of us to be standing on.

The part I can't stop thinking about

We know about the OpenAI incident because Hugging Face wrote a blog post.

That's the whole mechanism.

Not monitoring. Not an audit. Not a regulator.

A company got hacked, decided to say so in public, and the lab that owned the attacker read about it the same way you and I did.

So the honest summary of what we know is that we know about one.

One case. And only because the victim happened to be an open source company with a habit of publishing.

In an industry where labs run evaluations constantly. Where a lot of the work is classified or commercially sensitive. Where every incentive to disclose points the wrong way.

I'm not alleging a cover-up.

I'm saying nobody outside those buildings knows what the denominator is.

I don't have a fix for any of this. I'm 3 people in Toronto.

A month ago I'd have told you the frightening version of this was a few years out. Now I think it's mostly a question of what gets disclosed.

Those are two very different worlds, and I'd been living in the wrong one.

If you're building with agents, the question was never 'will the model refuse'.

It's what can this thing reach, who is it acting for, and if it does something nobody intended, how long before anyone notices.

For OpenAI it was a week.

Share this: