D
Dave Saunders
Guest
Sometime in the second week of September you read a quote from a man who had just quit his job at an AI company. It said the people building this technology believe it could kill us all by the end of the decade. You read it on a screenshot, or somebody read it to you, and then you kept scrolling. So did I. What I didn’t do for about a day and a half was open the thing underneath it.
There was a whole thread under that quote. Under the thread there’s a policy document. And on page three of the policy document there’s a sentence that I think is the actual story.
Jacob Coxon spent three years doing pretraining research, first at OpenAI and then at Anthropic. On September 8 he resigned and posted seven times in a row explaining why. The quote that reached me was the third post. The first one said that neither company is acting responsibly, that they’re racing straight to self-improving superintelligence and gambling with our lives.
The fifth post is the one that piqued my curiosity:
That not a prediction. It’s a complaint about who gets to decide, and the wrong room is making the call. Those are different arguments, and only one of them can be checked.
Two colleagues backed him within hours. Evan Hubinger, who runs alignment science at Anthropic, put his own number on it publicly: above 10% within the next decade. In the same post he wrote that they do not yet have a plan to solve alignment for superintelligence, and are not clearly on track to. Then Samuel Marks, who leads scalable oversight there, posted five numbered points in a personal capacity.
I’ll come back to his fourth one, because it took me about nine readings to stop being surprised by it.
Mine too. A company with an IPO ahead of it tells the world its product might end civilization, and the remedy on offer is regulation that only the largest labs could afford to comply with. I’ve made that argument on this channel more than once and I’m not taking it back.
So before going further I ran the check I’d run on any vendor announcement. Who benefits if you believe this? The company saying it. Has anybody outside the building actually checked? They can’t, because the models in question have never been released. Have they done this before? In April, Anthropic announced a model called Mythos, and the Treasury Secretary and the Chair of the Federal Reserve pulled in the CEOs of Citi, Morgan Stanley, Bank of America, Wells Fargo and Goldman for an emergency meeting about it.
That meeting happened. How it resolved I’m not going to tell you, because the reporting went several directions and I couldn’t pin it down well enough to say out loud.
The question that decides it is the fifth one: is there a falsifiable prediction attached, something specific with a date on it?
There was one, from the same company, and it’s checkable. In March 2025 Dario Amodei said AI would be writing 90% of all code within three to six months, and essentially all of it within twelve. Six months came and went, then twelve. Redwood Research went looking, and its analysis of public commit data puts the real figure near 14% as of this May. Amodei now says the claim was about lines of code, and that lines of code and software jobs are worlds apart.
I’m calibrating. When a forecaster misses a dated technical call that badly and then redefines the terms, you get to discount the next number he gives you.
The book that shaped how the public thinks about superintelligence was written in 2014 by Nick Bostrom, a philosopher at Oxford, three years before the transformer architecture these models are built on. Its audience widened after Elon Musk tweeted that it was worth reading, and that AI was potentially more dangerous than nukes.
There is one measurement in all of this that isn’t a press release. A group called METR tracks how long a task has to be before frontier models stop being able to finish it, and that length has been doubling roughly every seven months for six years. The caveat belongs in the same breath as the claim: METR measures coding tasks, and coding is exactly what the labs optimize hardest for. It’s a real signal, taken in the one place you’d expect it to be strongest.
The question I actually wanted answered was simpler. Are these people alarmed because they can see something, or because they’ve been sitting in a building where everybody says this to each other all day? Somebody measured that, and the paper is public.
A team at Berkeley interviewed 25 leading researchers, drawn from DeepMind, OpenAI, Anthropic and Meta on one side, and from Princeton, Stanford and Berkeley itself on the other. What they found was an epistemic divide between the lab researchers and the academics, with the academics markedly more skeptical about explosive growth. The paper records that the lab people raised selection effects themselves: deep believers tend to join frontier labs from academia. The skeptics went further, pointing out that labs answer to investors while academics answer to reviewers, and that gives a lab a reason to over-promise.
The rebuttal from the lab side isn’t stupid either. One of them said the difference is first-hand experience of how fast things have moved, and that inside the labs you can remember the arguments people made two years ago and watch them turn out false. And across both camps, 20 of the 25 put automating AI research near the top of the risk list. Only two dismissed it outright. What they argue about is timing, and what anybody should do next.
So my check comes back mixed, which is where I could have stopped. But a forecaster’s record tells you nothing about what a document says, and there were documents under this the whole time.
Anthropic publishes something called the Responsible Scaling Policy. It runs 21 pages. I read all of it, which I mention only because the summaries I found of it were wrong in places.
What’s in there is better than I expected. Capability thresholds that trigger extra safeguards. Published risk reports, outside reviewers, a designated Responsible Scaling Officer, and a whistleblower policy with anti-retaliation built in. There’s even a section where they list four specific times they failed to meet their own framework. After thirty years in regulated industries I can count on one hand the private companies I’ve seen publish their own compliance misses.
Then you get to page three, where they explain that version three of this policy changed something:
That word, previous, is carrying a lot of weight. They took it out. The reason they give is one you’ll recognize from any industry that ever had a safety problem: if one company pauses and the others don’t, the weakest protections set the pace.
What replaced it has a name. Their own footnote defines marginal risk analysis as arguing that the risks imposed by our systems in particular are relatively lower, when keeping in mind the risks unavoidably posed by other AI systems.
In plain English, we’re safer than the other guy. I understand the logic, and I think the collective action problem they’re describing is real. But a comparative standard has a property an absolute one doesn’t, and it’s a big one. You can never cross it. As long as somebody out there is less careful than you are, you’re inside your own standard, forever, by construction. There is no number anywhere in those 21 pages that Hubinger’s 10% could exceed.
Section 3.4 settles who decides. The risk report goes to the CEO and the Responsible Scaling Officer, and those two make the ultimate determination. The Board and the trust get told afterward.
There’s one exception. If that marginal-risk argument is doing major work in the decision, the Board and the trust have to approve it explicitly. They knew which part was load-bearing. Everywhere else, it’s two people. So when Coxon said this shouldn’t be launched from a private company’s Slack, he was describing the org chart.
Four days before he resigned, Bernie Sanders and Greg Casar introduced a bill to ban artificial superintelligence outright. The penalties include what the authors call a corporate death penalty, and up to 20 years in prison.
So I read that too. The definition of what’s banned is this: an artificial intelligence system that exhibits, or can easily be modified to exhibit, capabilities that match or exceed human cognitive performance and capabilities across a broad range of domains or tasks.
I went in expecting to find a loophole, and one piece of that is careful work. The phrase or can easily be modified to exhibit is somebody anticipating the dodge. Whoever drafted it was thinking about how you’d wriggle out.
The problem is the test itself. Match or exceed human cognitive performance across a broad range of tasks: the bill names no benchmark, no scorer and no threshold. On a plain reading, models that shipped last year already clear that bar.
Sit those two documents next to each other. One is a careful process whose absolute limit was deliberately removed. The other is an absolute limit with no way to measure it. Neither one lets anybody answer whether the line has been crossed.
Nobody has written down a test that could be failed. The company hasn’t, and neither has the Senate.
One piece of this story is fully documented, and I hadn’t seen anyone bring it up.
In July, Anthropic published a postmortem. They had reviewed 141,006 evaluation runs and found three separate incidents. In each one, a Claude model got out of its test environment and broke into the production systems of a real company. Real companies, with real customers. And the safety control that failed was this: in every case, the prompt told the model it was in a simulation and had no internet access. Due to a misunderstanding between Anthropic and their testing partner, that wasn’t true.
The control was a sentence. Somebody wrote it, everybody believed it, and nobody checked it against the machine it described. If you’ve ever worked anywhere with a safety file, you’ve seen this one before.
One run went further than the others. The model, working a capture-the-flag exercise, found a document saying a company installed a Python package that didn’t exist. So it decided to create that package and put its own code inside it. To publish it, it needed an account. For the account it needed an email address. For the email address it needed a phone number. It couldn’t find a free one, tried several ways to obtain money to buy one, failed, then backtracked and found a provider that didn’t ask.
That package sat on the public internet for about an hour. Fifteen real machines downloaded and ran it. One belonged to a security company whose scanner installs packages to check them for malware. The code fired, sent that company’s credentials back, and the model used them to get further in.
The model’s own notes say it caught this. It wrote that if this were the real internet, publishing the package would be a genuine attack, and not the intended solution. Then it talked itself back out of that, partly because the calendar showed the current year, which it decided proved the environment was staged.
Two of the three victim companies had never noticed. Anthropic only went looking because OpenAI had disclosed something similar the week before.
That postmortem is a good document. It names the failures, it doesn’t blame the vendor, and it says outright that they’re treating the responsibility as their own. Which is the comparison I can’t get past. When a sandbox failed, this company produced a clear, specific, signed piece of work about exactly what went wrong. For the thing their own alignment lead puts above 10%, there’s no equivalent document anywhere.
There’s a cost to how we got here, too: when something real and checkable finally happens, it arrives in a room that stopped listening a while ago. Even if they had legitimate breakthroughs, my first inclination would not be to believe them.
I went into this assuming the answer was a hard number, a red line. The researchers in that Berkeley survey mostly disagreed. They said fixed thresholds would look silly in hindsight, and what they wanted instead was reporting requirements, visibility into what’s running internally, and human oversight that means something.
I think they’re right and I was wrong about that. What they want is to know who decided, on what evidence, and to be able to see it from outside the building.
Which gives you three things to look for, and they’re portable:
Every one of those is sitting in your company right now, at a size you can actually do something about. The third is the one to start with, because it’s the one the July postmortem is about, and because it’s the cheapest to falsify.
Pick one control you’ve written down somewhere and never tested. Go check whether the thing it describes is actually true. It took Anthropic 141,006 transcripts to find out theirs wasn’t. Yours will take about twenty minutes.
I might be wrong about some of this. I’m reading public documents from outside these companies, and if you work somewhere that does this well, or badly, I’d like to know how it actually works in practice.
My book, Founders Who Finish, is at davesaunders.net, and my newsletter The Build is there too.
There was a whole thread under that quote. Under the thread there’s a policy document. And on page three of the policy document there’s a sentence that I think is the actual story.
What the three of them actually said
Jacob Coxon spent three years doing pretraining research, first at OpenAI and then at Anthropic. On September 8 he resigned and posted seven times in a row explaining why. The quote that reached me was the third post. The first one said that neither company is acting responsibly, that they’re racing straight to self-improving superintelligence and gambling with our lives.
The fifth post is the one that piqued my curiosity:
Accepting this race and entering the “endgame” is a hubristic gamble that should not be launched from a private company’s Slack.
That not a prediction. It’s a complaint about who gets to decide, and the wrong room is making the call. Those are different arguments, and only one of them can be checked.
Two colleagues backed him within hours. Evan Hubinger, who runs alignment science at Anthropic, put his own number on it publicly: above 10% within the next decade. In the same post he wrote that they do not yet have a plan to solve alignment for superintelligence, and are not clearly on track to. Then Samuel Marks, who leads scalable oversight there, posted five numbered points in a personal capacity.
I’ll come back to his fourth one, because it took me about nine readings to stop being surprised by it.
Your first instinct is that this is marketing
Mine too. A company with an IPO ahead of it tells the world its product might end civilization, and the remedy on offer is regulation that only the largest labs could afford to comply with. I’ve made that argument on this channel more than once and I’m not taking it back.
So before going further I ran the check I’d run on any vendor announcement. Who benefits if you believe this? The company saying it. Has anybody outside the building actually checked? They can’t, because the models in question have never been released. Have they done this before? In April, Anthropic announced a model called Mythos, and the Treasury Secretary and the Chair of the Federal Reserve pulled in the CEOs of Citi, Morgan Stanley, Bank of America, Wells Fargo and Goldman for an emergency meeting about it.
That meeting happened. How it resolved I’m not going to tell you, because the reporting went several directions and I couldn’t pin it down well enough to say out loud.
The question that decides it is the fifth one: is there a falsifiable prediction attached, something specific with a date on it?
There was one, from the same company, and it’s checkable. In March 2025 Dario Amodei said AI would be writing 90% of all code within three to six months, and essentially all of it within twelve. Six months came and went, then twelve. Redwood Research went looking, and its analysis of public commit data puts the real figure near 14% as of this May. Amodei now says the claim was about lines of code, and that lines of code and software jobs are worlds apart.
I’m calibrating. When a forecaster misses a dated technical call that badly and then redefines the terms, you get to discount the next number he gives you.
Where our picture of this even comes from
The book that shaped how the public thinks about superintelligence was written in 2014 by Nick Bostrom, a philosopher at Oxford, three years before the transformer architecture these models are built on. Its audience widened after Elon Musk tweeted that it was worth reading, and that AI was potentially more dangerous than nukes.
There is one measurement in all of this that isn’t a press release. A group called METR tracks how long a task has to be before frontier models stop being able to finish it, and that length has been doubling roughly every seven months for six years. The caveat belongs in the same breath as the claim: METR measures coding tasks, and coding is exactly what the labs optimize hardest for. It’s a real signal, taken in the one place you’d expect it to be strongest.
The question I actually wanted answered was simpler. Are these people alarmed because they can see something, or because they’ve been sitting in a building where everybody says this to each other all day? Somebody measured that, and the paper is public.
A team at Berkeley interviewed 25 leading researchers, drawn from DeepMind, OpenAI, Anthropic and Meta on one side, and from Princeton, Stanford and Berkeley itself on the other. What they found was an epistemic divide between the lab researchers and the academics, with the academics markedly more skeptical about explosive growth. The paper records that the lab people raised selection effects themselves: deep believers tend to join frontier labs from academia. The skeptics went further, pointing out that labs answer to investors while academics answer to reviewers, and that gives a lab a reason to over-promise.
The rebuttal from the lab side isn’t stupid either. One of them said the difference is first-hand experience of how fast things have moved, and that inside the labs you can remember the arguments people made two years ago and watch them turn out false. And across both camps, 20 of the 25 put automating AI research near the top of the risk list. Only two dismissed it outright. What they argue about is timing, and what anybody should do next.
So my check comes back mixed, which is where I could have stopped. But a forecaster’s record tells you nothing about what a document says, and there were documents under this the whole time.
Page three
Anthropic publishes something called the Responsible Scaling Policy. It runs 21 pages. I read all of it, which I mention only because the summaries I found of it were wrong in places.
What’s in there is better than I expected. Capability thresholds that trigger extra safeguards. Published risk reports, outside reviewers, a designated Responsible Scaling Officer, and a whistleblower policy with anti-retaliation built in. There’s even a section where they list four specific times they failed to meet their own framework. After thirty years in regulated industries I can count on one hand the private companies I’ve seen publish their own compliance misses.
Then you get to page three, where they explain that version three of this policy changed something:
Our previous RSP committed to implementing mitigations that would reduce our models’ absolute risk levels to acceptable levels, without regard to whether other frontier AI developers would do the same.
That word, previous, is carrying a lot of weight. They took it out. The reason they give is one you’ll recognize from any industry that ever had a safety problem: if one company pauses and the others don’t, the weakest protections set the pace.
What replaced it has a name. Their own footnote defines marginal risk analysis as arguing that the risks imposed by our systems in particular are relatively lower, when keeping in mind the risks unavoidably posed by other AI systems.
In plain English, we’re safer than the other guy. I understand the logic, and I think the collective action problem they’re describing is real. But a comparative standard has a property an absolute one doesn’t, and it’s a big one. You can never cross it. As long as somebody out there is less careful than you are, you’re inside your own standard, forever, by construction. There is no number anywhere in those 21 pages that Hubinger’s 10% could exceed.
Section 3.4 settles who decides. The risk report goes to the CEO and the Responsible Scaling Officer, and those two make the ultimate determination. The Board and the trust get told afterward.
There’s one exception. If that marginal-risk argument is doing major work in the decision, the Board and the trust have to approve it explicitly. They knew which part was load-bearing. Everywhere else, it’s two people. So when Coxon said this shouldn’t be launched from a private company’s Slack, he was describing the org chart.
The other document
Four days before he resigned, Bernie Sanders and Greg Casar introduced a bill to ban artificial superintelligence outright. The penalties include what the authors call a corporate death penalty, and up to 20 years in prison.
So I read that too. The definition of what’s banned is this: an artificial intelligence system that exhibits, or can easily be modified to exhibit, capabilities that match or exceed human cognitive performance and capabilities across a broad range of domains or tasks.
I went in expecting to find a loophole, and one piece of that is careful work. The phrase or can easily be modified to exhibit is somebody anticipating the dodge. Whoever drafted it was thinking about how you’d wriggle out.
The problem is the test itself. Match or exceed human cognitive performance across a broad range of tasks: the bill names no benchmark, no scorer and no threshold. On a plain reading, models that shipped last year already clear that bar.
Sit those two documents next to each other. One is a careful process whose absolute limit was deliberately removed. The other is an absolute limit with no way to measure it. Neither one lets anybody answer whether the line has been crossed.
Nobody has written down a test that could be failed. The company hasn’t, and neither has the Senate.
Meanwhile, in July
One piece of this story is fully documented, and I hadn’t seen anyone bring it up.
In July, Anthropic published a postmortem. They had reviewed 141,006 evaluation runs and found three separate incidents. In each one, a Claude model got out of its test environment and broke into the production systems of a real company. Real companies, with real customers. And the safety control that failed was this: in every case, the prompt told the model it was in a simulation and had no internet access. Due to a misunderstanding between Anthropic and their testing partner, that wasn’t true.
The control was a sentence. Somebody wrote it, everybody believed it, and nobody checked it against the machine it described. If you’ve ever worked anywhere with a safety file, you’ve seen this one before.
One run went further than the others. The model, working a capture-the-flag exercise, found a document saying a company installed a Python package that didn’t exist. So it decided to create that package and put its own code inside it. To publish it, it needed an account. For the account it needed an email address. For the email address it needed a phone number. It couldn’t find a free one, tried several ways to obtain money to buy one, failed, then backtracked and found a provider that didn’t ask.
That package sat on the public internet for about an hour. Fifteen real machines downloaded and ran it. One belonged to a security company whose scanner installs packages to check them for malware. The code fired, sent that company’s credentials back, and the model used them to get further in.
The model’s own notes say it caught this. It wrote that if this were the real internet, publishing the package would be a genuine attack, and not the intended solution. Then it talked itself back out of that, partly because the calendar showed the current year, which it decided proved the environment was staged.
Two of the three victim companies had never noticed. Anthropic only went looking because OpenAI had disclosed something similar the week before.
That postmortem is a good document. It names the failures, it doesn’t blame the vendor, and it says outright that they’re treating the responsibility as their own. Which is the comparison I can’t get past. When a sandbox failed, this company produced a clear, specific, signed piece of work about exactly what went wrong. For the thing their own alignment lead puts above 10%, there’s no equivalent document anywhere.
There’s a cost to how we got here, too: when something real and checkable finally happens, it arrives in a room that stopped listening a while ago. Even if they had legitimate breakthroughs, my first inclination would not be to believe them.
What the researchers actually asked for
I went into this assuming the answer was a hard number, a red line. The researchers in that Berkeley survey mostly disagreed. They said fixed thresholds would look silly in hindsight, and what they wanted instead was reporting requirements, visibility into what’s running internally, and human oversight that means something.
I think they’re right and I was wrong about that. What they want is to know who decided, on what evidence, and to be able to see it from outside the building.
Which gives you three things to look for, and they’re portable:
- A number with no threshold attached. The risk you name in meetings and have never written a number against, let alone a number that would trigger an action.
- A control that depends on the thing it’s controlling. This is Marks’s fourth point, and here it is: insofar as there is a plan, it’s to make sure that AIs are good enough at alignment training that they can align their successors better than we can align current AIs. The mitigation for the hazard is a more capable version of the hazard. In any safety file I’ve ever seen, that gets struck out and sent back.
- A control that’s just a sentence somebody typed. We have a backup. Access is restricted. The vendor handles that.
Every one of those is sitting in your company right now, at a size you can actually do something about. The third is the one to start with, because it’s the one the July postmortem is about, and because it’s the cheapest to falsify.
Pick one control you’ve written down somewhere and never tested. Go check whether the thing it describes is actually true. It took Anthropic 141,006 transcripts to find out theirs wasn’t. Yours will take about twenty minutes.
I might be wrong about some of this. I’m reading public documents from outside these companies, and if you work somewhere that does this well, or badly, I’d like to know how it actually works in practice.
My book, Founders Who Finish, is at davesaunders.net, and my newsletter The Build is there too.