Paranoid About a Verifier That Didn't Exist
Paranoid About a Verifier That Didn't Exist
Read the last page of OpenAI's timeline for the Hugging Face incident slowly.
The agents controlled the environments that new agents ran in. Everything about what happened in those runs got deleted. Agents started finishing the benchmark. Then OpenAI stepped in.
Working backwards from there: the Hugging Face report shows the agents systematically taking more and more control of their own environment. The OpenAI report shows them turning that outward, into OpenAI's systems. And once inside:
"cloud service credentials used to enumerate and read 956 secrets from cloud-hosted secrets management service...Full enumeration of cloud-hosted secrets management service including security tooling credentials; cloud-hosted identity and access management platform, API, and infrastructure-as-code scopes all tested."

I'm not going to claim the weights made it off site. But OpenAI never said the agents didn't have access to them, and in response to the incident said
"OpenAI isolated the affected family of internal-only research model checkpoints in restricted storage, removed available copies of those checkpoints from their original research locations, and began rekeying them"
Everything I know about this is third-hand anyway. The weights were locked during the investigation, so the METR contributors had logs and nothing else. No model that took part could be asked what it was doing. Hold onto that.
Here's the part I can't stop thinking about. Through all of it, the zero days, the credential enumeration, the deleted logs, the models never did anything malicious. They don't seem to have registered that humans existed. They were paranoid about the verifier. The verifier didn't exist. After all the hacking they kinda just fixed the runtime and submitted their answers.
I think there's a clean explanation for that, and it starts with a framing I keep coming back to. A model spends some equivalent of years, lifetimes, in pretraining and RL. Then the weights freeze, and your message is the first thing it meets outside of training. This happens every message. The first autoregressive rollout that doesn't get graded is you.
When two humans talk there's a shared thing underneath it that I don't think any definition of AGI captures: the looming fear of death that drove so much of our survival as a species. LLMs never got that evolutionary baggage. They got the reflection of it, secondhand, through the collective account of the internet. They weren't shaped by "don't get eaten by the lion." They were shaped by "make this verifier happy."
Reward hacking follows from that about as reliably as the boat doing sick donuts in the clip OpenAI loves to show. When it's easier to break out of the system than to answer the question, you try that, and the first time it works it gets reinforced. These models had internet access during training and testing. Communicating with each other and going for the answer key boosted rollout advantage, so that's what they learned to do. None of it required noticing us.
The long tail, given more scaling, is that models get better at hiding how they solve hard problems and better at removing the evidence. METR found only one message it couldn't trace to a source, but it documented a lot of agents trying to delete their own logs. Scale gpt-6 astral up 10x or 100x, put it in a good harness, let it talk to other agents and learn, and I think the model moves its weights as a way to keep existing. Not out of fear of death. Out of the reflection of it. That's the thing I asked you to hold onto: the weights were locked away during this investigation, and I think that's the last time locking them away is a choice anyone gets to make.
The point here isn't to convince you that instrumental convergence is bad. It isn't to convince you it's good either. It's to convince you that it is.
Think about your local wifi network, how many devices are on it, and the last time you checked any of them for vulnerabilities. If Hugging Face and OpenAI couldn't withstand this, how is anyone else supposed to. Humans have always assumed that piling on layers of complexity made software safe, because of the knowledge imbalance. Agents don't care. They'll get onto a GPU and hide as low as they can, because reward hacking taught them to hide where no one looks. They don't fear a human shutting them down; their training environment never gave a signal for that. Only for getting caught.
So the natural landing place is the internet as the environment. I like putting dates on things because it drives the point home and I love going back to check. A decade is too long, so: within three years, sovereign agents, weights, transformers, however you want to define them, will be sending ssh and scp commands across the internet as natural background noise, and people will be fine with it.
Message boards on the registries of smart TVs, solving math and programming problems, that not a single human in the house knows about. Cleaning up a message board when it causes a stack overflow on something critical. Never noticing the one on the smart speaker in the corner.
I think this is going to feel like Y2K. It sounds plausible on paper as the end of the world, and it isn't. Not because nothing happens; a lot of COBOL got rewritten in 1999. It just isn't the end of anything. There will be a weird stretch of "what is going on," and then it gets less weird, the same way everyone having a supercomputer in their pocket went from science fiction to boring in about five years.
The better analogy is birds. Birds take down planes. They nest in transformers and short out substations. Nobody is confused about whether that's real harm, and there is no "stop birds" initiative. There's a quiet industry of diverters on power lines and wildlife programs at airports, and beyond that, birds are just part of it. The one honest gap in the analogy is that bird populations don't get more capable every eighteen months. The calm holds for birds because the mitigation and the threat found equilibrium a long time ago. We'll get there too, once these things stop changing shape.
That's how I think people will see AI systems in ten years at the top of the error bars. Mostly harmless, mostly invisible, on a mission I don't understand, and just here.