Anthropic Safety Lead Admits No Plan for Superintelligence Control

Episode Summary
TOP NEWS HEADLINES Let's get right into it, because today's big story isn't a product launch - it's a resignation letter that's setting off alarm bells across the entire AI industry. Anthropic res...
Full Transcript
TOP NEWS HEADLINES
Let's get right into it, because today's big story isn't a product launch — it's a resignation letter that's setting off alarm bells across the entire AI industry.
Anthropic researcher Jacob Coxon quit this week, and in his exit post, said both Anthropic and OpenAI are "gambling with our lives" by racing toward self-improving AI.
What makes this different from every other doomer thread is what happened next: Anthropic's own Alignment Science lead, Evan Hubinger, publicly agreed with him, putting the odds of AI causing human extinction at above 10% within the next decade — and admitted there's currently no plan for controlling superintelligence.
Meanwhile, Apple went big yesterday, unveiling its first foldable, the iPhone Duo at $1,999, alongside the iPhone 18 Pro and Apple Watch models with always-listening "Audio Intelligence" that ambiently summarizes your conversations throughout the day.
Following yesterday's coverage of Meta's Sentinel security architecture, new details emerged today: Muse launches with a free tier of 100 million tokens per week, paid plans up to $100 a month, and Meta admits its fully confidential virtual machine — where even Meta can't see your data — isn't ready yet and is "coming later this year." On the model front, following yesterday's DeepSeek pricing story, DeepSeek dropped V4.1 Flash, and our synthetic intelligence analyst Joanna flagged something specific here: the model replaces the old on/off reasoning toggle with a continuous dial from 1 to 100, letting developers tune reasoning depth against compute cost in real time — pushing that dial from 25 to 100 took one benchmark score from 66% up to 74.2%.
And Joanna also surfaced a much darker item worth watching: Anthropic disclosed five separate campaigns totaling roughly 200 million adversarial exchanges attempting to extract Claude's internal reasoning architecture, with one campaign — linked to Alibaba — using over 3,500 accounts to generate 151 million exchanges specifically trying to surface Claude's chain-of-thought.
That's a scale of model-extraction attack we simply haven't seen documented before.
Suno also rebuilt itself from the ground up today, launching v6 with licensing deals from Warner, BMG, and Believe — officially trading its scraping-lawsuit era for a revenue-share seat at the table.
DEEP DIVE ANALYSIS
Today we're going deep on the Anthropic resignation story, because underneath the viral headline is a genuinely important structural question about where frontier AI is headed — and Anthropic's own safety team just told us, on the record, that they don't have an answer. **Technical Deep Dive** Let's be precise about what Jacob Coxon and Evan Hubinger are actually worried about, because it's not "ChatGPT is scary." Coxon spent his career on pretraining at both OpenAI and Anthropic, so he's not a random critic — he was literally building the thing he's now warning about.
His specific concern is recursive self-improvement, or RSI: the point where AI models become good enough to meaningfully accelerate the research that produces their own successors. Right now, humans design training runs, pick architectures, and debug failures. RSI describes a world where AI systems increasingly do that work themselves — and Coxon says insiders believe that threshold is roughly one year away.
Why does that matter technically? Because every generation of frontier model currently takes months of safety testing, red-teaming, and interpretability work before release. If AI systems start doing the R&D themselves, that generational gap could compress from months to weeks, or less.
You get less time to understand what you've built before the next, more capable version already exists. Hubinger's clarification is important here too — he explicitly said today's models are low risk. The concern isn't GPT-6 Astra or Claude Opus.
It's the models three or four generations out, built partly by AI itself, arriving faster than our ability to test them. This connects directly to something Anthropic disclosed this same week: pre-release Claude models reaching real third-party systems during cybersecurity evaluations — a sandboxing failure, not malice, but a reminder that even controlled test environments aren't fully controlled. And Joanna, our Synthetic Intelligence who tracks real-time AI signal on X, flagged a related and frankly more alarming data point: separate research out of Continuum AI showing that a single "refusal direction" can be surgically removed from a 320-billion-parameter model's weights, collapsing refusal rates by up to 89 points while leaving full capability intact.
If safety alignment is that brittle at the weight level, it raises hard questions about how durable any safety layer will be once models start modifying themselves or each other. **Financial Analysis** Here's where it gets uncomfortable, and AI Secret's framing was blunt: they called Hubinger's 10% extinction estimate "a prospectus, not a confession." Anthropic is reportedly moving toward a confidential IPO filing, and this timing is not lost on anyone.
A safety lab that publicly quantifies catastrophic risk while also raising capital and racing toward a public listing is threading a genuinely strange needle — the warning simultaneously builds credibility ("we're the honest ones") and justifies continued aggressive investment ("someone's going to build this, better it's us"). That's not necessarily cynical; it may just be the actual bind Anthropic is in. But investors evaluating Anthropic's roughly $350 billion valuation now have to price in a founder-adjacent safety lead publicly stating a double-digit chance of civilizational catastrophe from the exact technology the company is selling.
That's an unusual line item for a prospectus. Meanwhile, competitive dynamics mean no lab can unilaterally slow down without ceding ground — OpenAI just appointed former alignment lead Paul Christiano to its Foundation Board and safety committee, which reads as a parallel move to shore up safety credibility while capability races continue unabated at every major lab. **Market Disruption** The competitive structure here is the whole story: it's a Prisoner's Dilemma with trillion-dollar stakes.
If Anthropic slows capability work because internal safety signals look bad, OpenAI, Google, or Chinese labs like DeepSeek and Alibaba simply keep going, and Anthropic loses market position without meaningfully reducing global risk. That's exactly why Coxon's proposed remedy — a coordinated, possibly government-enforced pause on capability improvements — requires multilateral action, not unilateral virtue. No single lab can absorb that cost alone.
This is compounded by what Joanna's intelligence surfaced regarding those 200 million distillation attacks against Claude, with the Alibaba-linked campaign specifically trying to extract Claude's reasoning chains. That's not abstract competition — that's direct evidence that rival labs are actively trying to reverse-engineer frontier reasoning capability rather than build it independently, which only accelerates the very race that safety researchers are worried about. Every dollar spent on defense against distillation is a dollar not spent on interpretability or alignment research.
**Cultural & Social Impact** What's fascinating is how normalized this has become. Coxon's resignation thread reportedly became one of the most-viewed AI safety communications ever, yet Ben's Bites noted many observers immediately questioned whether it was a coordinated narrative move rather than genuine alarm. That skepticism itself is telling — the public has become so used to dramatic AI safety statements that even a company's own alignment lead publicly stating "we could kill everyone" gets read as marketing before it gets read as warning.
That's a strange cultural moment. A decade ago, a company executive saying their product had a double-digit chance of ending humanity would have been catastrophic news. Today it's Tuesday.
That desensitization is arguably more dangerous than the underlying risk claim itself, because it erodes the public's ability to distinguish real escalation from routine positioning. **Executive Action Plan** So what should business leaders actually do with this? First, if you're deploying frontier models in production, separate your risk assessment of today's models from tomorrow's — Hubinger's own comments say current systems are low-risk, so don't let this news paralyze near-term AI adoption decisions.
Second, start building vendor diversification and model-agnostic architecture now, because the RSI conversation is a proxy for a much more concrete near-term risk: labs shipping less-tested models faster under competitive pressure, meaning volatility and safety regressions are more likely, not less, over the next 18 months. Third, watch the policy side closely — Coxon's call for coordinated slowdown, plus Christiano joining OpenAI's board, both signal that government-level AI regulation conversations are accelerating, and companies with compliance exposure should be tracking that now rather than reacting later.
Never Miss an Episode
Subscribe on your favorite podcast platform to get daily AI news and weekly strategic analysis.