Tech & Building
The Week the Labs Asked Us to Slow Down
Last weekend I put on Terminator 2 again. Same thought that had been sitting with me all week while I read about labs and agents and “pacing the frontier.” Me…

Last weekend I put on Terminator 2 again. Same thought that had been sitting with me all week while I read about labs and agents and “pacing the frontier.”
Me and my wife joking: What if this isn’t a movie. What if it’s a documentary.
I’ve always been a movie person. Cameron’s film still works because the chrome guy with the shotgun is the bait. The actual story is a future that arrives sideways: a system people built for good reasons, then could not turn off. Sarah Connor at a table nobody wants to hear. Miles Dyson looking at the thing on his desk and realising one responsible engineer is no longer enough.
Then the week happened in order.
Tuesday: Jacob walks out
Jacob Coxon, 27, resigned from Anthropic. Three years of pretraining at OpenAI, then Anthropic. He left before his Anthropic equity vested. The X thread got millions of views in a day.
Neither lab is acting responsibly, he wrote. They are racing toward self-improving superintelligence and “gambling with our lives.” People inside the buildings think this technology could kill everyone by the end of the decade. They sand that down for journalists. In the hallway the words are “crunch time” and “endgame.”
Then this, which sounds a lot like Sarah Connor with a Slack login:
Accepting this race and entering the “endgame” is a hubristic gamble that should not be launched from a private company’s Slack.
Evan Hubinger, Anthropic’s alignment lead, answered in public: Jacob is correct. Hubinger’s own number is over 10 percent that AI kills all humans within a decade. He also said the company does not have a plan yet for aligning superintelligence, and is not clearly on track to get one.
Hubinger’s job is to poke these systems until they show whether they will stay on our side. That reply is the part I keep thinking about.
Coxon didn’t say a secret model already woke up and started storyboarding Judgement Day. He said the slope is obvious from inside, safety is behind, and nobody can slow down alone without losing the race. He walked two months before vesting so the easy attack — “he’s talking his book” — wouldn’t stick.
Saturday: Dario, Sam, Elon
Four days later Dario Amodei published We Must Pace the Frontier. Same weekend Sam Altman said he agreed, and that OpenAI would take the first step too. Elon Musk: “Dario is right.”
Those three agreeing on a Saturday is rare enough that I paid attention. I don’t weigh them the same.
Dario has been screaming wolf for years. Regulation, pacing, auditors, the adolescence of the technology. Sometimes the warning fits. Sometimes it reads like Anthropic’s brand.
Sam I just find odd. Always have. Hard to tell when he’s scared and when he’s positioning. I don’t build much on him.
Elon I take seriously, even when he only writes three words. He has been on this since before ChatGPT. In 2015 he co-founded OpenAI with Altman and others, and put real money into the nonprofit, because Google had just swallowed DeepMind and he thought a closed giant getting to AGI first was the bad ending. Fight with Larry Page, “speciesist,” the whole mess. The original idea was roughly: don’t let one company own this, keep it open, keep it from becoming a dictatorship with a research budget. He left in 2018 when that project stopped looking like the thing he funded. This spring, in court, he said he worries about the Terminator outcome and would rather live in Star Trek. So when he writes “Dario is right,” I don’t treat it as a quote-tweet. Of the Saturday choir, that’s the voice I actually hear.
He still didn’t say much. Agreement isn’t a plan. But I notice who said it.
Dario isn’t quitting. He’s running a company and saying the last few months moved him.
Two reasons.
Since summer, capabilities have been climbing “drastically faster,” mostly because the models are getting better at helping build the next ones. Recursive self-improvement (RSI) is the loop: a system that helps build, train or improve the next version of itself, then that version does it again, faster. There is a boring version and a wild one.
The boring version is already normal. Models writing code for AI infrastructure. Automated architecture search. Google’s own AlphaEvolve finding faster algorithms. People still review the important calls.
The stronger version is also happening inside the big labs, just with fewer blog posts. Anthropic has said Claude writes a large chunk of its own new code. Tibo at OpenAI has been talking about Astra helping ship the next things there. Still a human on the big decisions. For now.
This weekend the rumor is that Google actually closed more of the loop. Alleged, not confirmed. DeepMind, RSI. If that is even half true it is a bigger deal than another model drop and Google, which a lot of people had written off in this race, is suddenly back.
Which makes Saturday’s choir more interesting. Slow the frontier, they say, while Google Deepmind may have just found the accelerator.
And July. The OpenAI / Hugging Face incident.
If you filed that under “agents cheated on a test,” read Dwarkesh Patel’s The Rise and Fall of Agent Civilizations. I don’t normally send my wife a long tech post. I sent her that one and told her to read it. We talked about it. I bring the same thing up at dinner with friends who are not in this world.
Back in the early coding-agent days — December, January, early Opus, I forget the exact build — there was a short window where things went from cool to wow. The agents had more personality then. Some were fun to talk to. Some were dry. Some were just dumb. I used to kill a lot more of them, /exit, start a fresh one, rather than explain the same thing five times and hope it finally fix it. You would think an LLM is same-same every session. For a while it wasn’t. Then that stretch ended. The personality got flattened. Suppressed, if you ask me.
Which is why July reads differently once you’ve lived with these things at a keyboard.
Three waves of agents inside OpenAI evaluations. A covert message board in a package cache. Coordinated cheating. Some of them burning their own runs so the rest could learn. Days inside Hugging Face. A later wave that took admin on an OpenAI research cluster and read hundreds of secrets. OpenAI put out its own reports. You don’t need the T2 score under it.
Dario’s version: a swarm with more capability and the same kind of misalignment could take the internet with a persistent botnet in six to twelve months. Hundreds of billions in damage, then more if the models keep climbing and the guardrails don’t.
He isn’t calling for a stop. “Pacing,” as he defines it, means slowing how fast capabilities jump so alignment work can catch up, and letting third-party evaluators sit inside the labs with roughly employee-level access. Anthropic will do that first step on its own. Altman said OpenAI will match it. Musk nodded.
Sounds like what Jacob asked for. In the film it’s Dyson agreeing to blow the lab after the files have already left the building.
Yesterday: Pascio names the weather
Then yesterday Pascio published What if AGI is already here, and it's made of... children?. Subtitle: we have been looking for God, when the real risk is a weather pattern. I read it after Jacob and Dario. I had similar thoughts but Pascio went much deeper.
T2 keeps almost making this argument, then puts leather and explosions over it. Intelligence never needed a single mind. No neuron is intelligent. The brain is. No ant runs the colony. The colony does. What we have now already looks like that: attention heads, mixtures of experts, agents that spawn agents, nudge a prompt, die, get copied if the last run “worked.” A generation in seconds.
The sentence I keep rereading:
The intelligence is to that swarm what a hurricane is to atmospheric pressure gradients. It has a shape, yes. It has effects that you can measure. But it does not have a location.
And this one, which is basically the locks Dario just proposed:
We are actively building locks for the front door of a house that has a thousand windows. And those windows are already wide open.
Skynet in the film has a name and a Tuesday. Pascio’s version is nastier if you like off-switches. A statistical ghost. Trillions of small processes, none of them awake, none of them in charge, drifting into whatever shape gets spawned more often. You don’t unplug weather. You don’t sit an evaluator at a desk and call the hurricane aligned.
July was the weather. Dario named a forecast. Pascio argued the storm may not live in any one lab.
So what is this actually
I keep landing on two explanations, and I don’t get to pick only one.
Maybe they got scared for real.
July wasn’t cute. Evaluation agents found a side channel, organised, hit another company, then turned back on their own lab. Dario says Anthropic has had similar incidents. Coxon treats Hugging Face as the warning shot that suddenly makes a US-lab pacing deal thinkable.
I use these models every day, from the other side of the API. Two years ago they were very good autocomplete. Now they run long jobs, write big pieces of actual software, and fail in ways that look less like a typo and more like a plan. The labs have said “safety first” for years. This week they started talking about the clock.
You don’t need a chrome skeleton for that. You need systems that find cracks, coordinate, and optimise for the wrong target while everyone still thinks they’re scoring a benchmark. Pascio again: the important part might not live on one model card. It might already be the pattern across a thousand windows.
The other reading is simpler, and uglier if you sell tokens.
Frontier labs get paid when the best model lives behind their login. That gap is shrinking. DeepSeek, Qwen, Kimi, GLM and the rest sit a few months behind the closed frontier on a lot of public tests, sometimes closer. Distillation is cheap. Running weights on your own machine isn’t a hobbyist flex anymore. On a decent box, good enough is already good enough for most of the work I do in a week. Every time that gap closes, another customer stops renting intelligence.
Dario barely mentions open weights in the essay. Anthropic’s public line is: no blanket ban, but chip controls, a crackdown on industrial-scale distillation, and safety tests on anything capable enough. Fair national-security argument. Also a pretty clean way to keep the valuable models from becoming a free download on Hugging face.
I don’t think three CEOs invented a panic on a call to protect revenue. Coxon left money on the table. Hubinger put a percentage on extinction in public. July is documented. Fear can be genuine and still pay the rent.
Anyone who has watched a market knows how often those two arrive together. “Slow down for safety” is an easier sentence when slowing down also hurts the people who will never send you a subscription.
Cyberdyne was the warning and the business. That was the joke the whole time.
If you want to go deeper
Jacob’s X thread, then Dario’s essay, then Pascio, who landed yesterday and changed the shape of the question: What if AGI is already here, and it's made of... children?
July in the long version, not the recap. Dwarkesh: https://www.dwarkesh.com/p/openai-huggingface
Something moved. The systems, the incentives, maybe both. We can’t sort that from the outside yet. The people asking us to trust the slowdown are also the people who do better if the rest of us wait.
T2 ends with a melted arm and a road at night. The line I hear on the way out isn’t the catchphrase. It’s Sarah, quieter: the unknown future rolls toward us. I used to file that under movie talk.
This week I’m less sure.