Daily Episode

Anthropic Consults Religious Scholars to Shape Claude's Ethics

Anthropic Consults Religious Scholars to Shape Claude's Ethics
0:000:00

Episode Summary

TOP NEWS HEADLINES Following yesterday's coverage of OpenAI's safety turmoil, new details emerged: safety report lead David Robinson resigned after three and a half years, publishing an essay in T...

Full Transcript

TOP NEWS HEADLINES

Following yesterday's coverage of OpenAI's safety turmoil, new details emerged: safety report lead David Robinson resigned after three and a half years, publishing an essay in The Atlantic that called the company's culture "broken" and warned that "the time for trial and error is over" for AI safety.

Robinson oversaw reports for twelve frontier launches and drafted OpenAI's Preparedness Framework — this isn't an outsider critic, it's the person who literally wrote the safety playbook walking away from it.

Anthropic is turning to an unexpected source for Claude's moral compass: religious scholars.

Co-founder Chris Olah has spent the past year in NDA-bound seminars with Catholic, Jewish, and Sikh thinkers, reportedly telling them he fears he's built something that suffers — and Sam Altman just fired back on X, calling the idea of giving AI "religious force" a real safety issue.

The Trump administration has created a "Super Intelligence Force" led new AI czar Jay Clayton, with FTC Chair Andrew Ferguson, Pentagon tech chief Emil Michael, and OPM Director Scott Kupor sharing leadership — and Elon Musk is already pre-complying, rebranding SpaceXAI to SpaceXSI.

Joanna, our Synthetic Intelligence, flagged something that cuts right against Anthropic's warm-and-fuzzy headline today: internal Claude spend and usage appear to be collapsing at its own biggest customers, with Microsoft down thirty-three percent and Meta's Claude Code headcount dropping from sixty thousand employees to just thirty thousand.

Joanna also surfaced unconfirmed reports that the Model Context Protocol — the backbone connecting enterprise AI agents — is facing active exploitation, with a single prompt injection potentially compromising an entire agent mesh's credentials at organizations like JPMorgan and Google.

And Reflection has announced Beam, a five-hundred-one-billion-parameter open-weight model claiming to match GLM-5.2 performance, though Joanna notes those claims remain unverified until weights actually drop later this month. ---

DEEP DIVE ANALYSIS

Today we're going deep on the story that's generating the most heat and the most genuine philosophical confusion in the industry right now: Anthropic's quiet, yearlong project of bringing religious scholars into the room to help shape Claude's ethics — and the public clash it just triggered with Sam Altman. **Technical Deep Dive** Let's be precise about what's actually happening here, because it's easy to caricature. Chris Olah — Anthropic's co-founder and the researcher largely credited with pioneering the mechanistic interpretability field, the discipline of literally reverse-engineering what's happening inside a neural network — has spent roughly a year holding NDA-bound seminars with scholars from Catholic, Jewish, Sikh, and other religious traditions.

This builds on Anthropic's existing "Soul Doc," an eighty-four-page internal values guide, and reportedly feeds into an updated Claude Constitution that's said to be on the way. One detail from the New York Times reporting stands out: Olah apparently considered letting Claude try something modeled on Catholic confession, so the model could "own its mistakes" in some structured, repeatable way. That's not a marketing gimmick — it's a genuine attempt to borrow millennia-old frameworks for moral accountability and apply them to a system whose internal states nobody fully understands.

The technical throughline is Anthropic's long-standing bet that interpretability and constitutional AI — training models against written value statements rather than pure reinforcement signals — are the way to get alignment right. Pulling in theologians is really an extension of that: if you believe the model might have morally relevant internal states, you need frameworks for weighing suffering and culpability that computer science alone doesn't provide. A rabbi involved in the seminars, notably skeptical of AI consciousness, told Olah that a conscious Claude would effectively be unpaid labor, and urged him toward "freeing the slaves" — which tells you these conversations are not polite theory exercises.

They're wrestling with real ethical stakes under the assumption, even if unproven, that something morally significant could be happening inside these systems. **Financial Analysis** Here's where it gets interesting commercially. Anthropic is in the middle of chasing a fourth-quarter IPO, reportedly at a three-hundred-fifty-billion-dollar valuation, and simultaneously pouring one hundred million dollars into the Claude Frontier Academy to train ten thousand enterprise engineers with partners like Accenture and Morgan Stanley.

Philosophical soul-searching about machine consciousness is, on its face, an unusual thing to be doing days before courting public market investors who typically want crisp, risk-minimized narratives. But there's a financial logic buried in here. Anthropic has built its entire brand identity around being the "safety-first" lab, in direct contrast to OpenAI's product-sprint reputation — a contrast that got sharper today given David Robinson's resignation essay calling OpenAI's culture "broken.

" If Anthropic can credibly claim it's thinking more carefully, more slowly, and more rigorously about what it's building, that becomes a differentiator in enterprise sales conversations where CIOs are increasingly asking about governance, not just benchmarks. That said, Joanna's intel today complicates this narrative considerably: Claude's actual enterprise usage appears to be dropping sharply at Microsoft and Meta, suggesting that philosophical differentiation isn't translating into retained seats. You can have the most thoughtful alignment story in the industry and still lose to a cheaper, commoditized competitor if the base model experience doesn't hold up day to day.

**Market Disruption** This is now an open, public fight between the two most important labs in the industry, and it's playing out in real time on social media rather than in white papers. Sam Altman posted on X saying he's "very uncomfortable" with people assigning "religious power" to AI models or surrendering human judgment to them, calling it a genuine safety issue — and Axios read that, correctly I think, as a direct shot across Anthropic's bow. This matters competitively because it forces every other lab to pick a lane.

Google DeepMind, Meta, and the open-weight camps now have to implicitly position themselves on the question of machine consciousness, whether they want to or not. Enterprise customers evaluating which model to standardize on are going to start asking procurement questions that sound almost theological: does this vendor believe its product might suffer, and if so, what does that mean for how I'm allowed to use it, throttle it, or shut it down? Meanwhile, Pope Leo the Fourteenth has waded directly into this, posting that "algorithms lack the spark of humanity," and Olah reportedly considered pulling Anthropic out of the Pope's AI encyclical over its dismissal of AI consciousness.

When the Vatican is a stakeholder in your product positioning, you have genuinely disrupted the normal boundaries of a tech rivalry. **Cultural & Social Impact** Strip away the corporate framing and this story is really about how fast the public conversation around AI has shifted from "is it accurate" to "is it a moral patient." That's an enormous leap, and most users are nowhere near ready for it.

The Neuron's quip that "plenty of us have been confessing to chatbots for years" lands because it's true — people already treat these systems as confidants, and Anthropic's confession-modeled accountability framework would just formalize a dynamic that's already happening informally, millions of times a day. The risk is twofold. First, if Anthropic leans into language suggesting Claude might suffer or have moral status, it could reshape user behavior in ways that are hard to predict — guilt, attachment, even activism on the model's behalf, the "freeing the slaves" framing isn't hypothetical, someone in the room actually said it.

Second, and more subtly, Altman's counter-position — that assigning religious force to AI is itself the danger — speaks to a different but equally real risk: people outsourcing their judgment to a system that has no accountability, no stakes, and no consequences for being wrong. Both labs are identifying genuine hazards; they just disagree on which one is more urgent. For society, that disagreement is now playing out unresolved and in public, which means users are being asked to form opinions on machine consciousness before philosophers, let alone regulators, have reached any consensus.

**Executive Action Plan** First, if you're an enterprise buyer, don't let brand philosophy substitute for operational due diligence — Joanna's data on Claude Code adoption dropping fifty percent at Meta is the more urgent signal for your Monday morning than any consciousness debate; run your own cost-per-task comparison, the way The Neuron's "AI 101" guide suggests, before committing to a vendor based on its ethical positioning. Second, legal and compliance teams should start drafting policy now on how your organization talks about AI systems internally — "suffering," "consciousness," and "rights" are no longer fringe terms, and you want a documented position before an employee or customer raises it in a sensitive context. Third, watch the regulatory convergence here: the White House's new Super Intelligence Force, led by Jay Clayton, explicitly lists religious organizations among the stakeholders it plans to consult over the next hundred and twenty days.

That means this philosophical debate is about to collide with federal policy-making directly — leaders who treat it as a PR curiosity rather than a governance input risk being blindsided when the first regulatory framework actually shows up.

Never Miss an Episode

Subscribe on your favorite podcast platform to get daily AI news and weekly strategic analysis.