Special Episode

No Skin in the Game: The One Kind of Reasoning Your AI Can't Do

No Skin in the Game: The One Kind of Reasoning Your AI Can't Do
0:000:00

Episode Summary

Lia: In February this year, Demis Hassabis stood on a stage in Bangalore and proposed a test for artificial general intelligence. Not a benchmark. Not an exam

Full Transcript

Lia: In February this year, Demis Hassabis stood on a stage in Bangalore and proposed a test for artificial general intelligence. Not a benchmark. Not an exam. A test. Thom: Train an AI on everything humanity knew up to nineteen-eleven. Then see if it discovers general relativity. Lia: His own verdict on whether today's systems could pass it was one sentence long. It's clear they couldn't. Thom: And here's the part I keep chewing on. Three weeks *before* he said that, a researcher inside his own lab had already published the paper explaining why. Lia: Today we're going to walk you through that argument — and then we're going to show you where it broke, four days ago, in a way that turns a philosophy paper into something you can actually use tomorrow morning. Thom: Because the interesting question isn't whether your AI can be Einstein. It's whether your AI has anything at stake when it's wrong. Lia: Spoiler: it doesn't. And that has consequences for how you're deploying it right now. Lia: So let me set a scene. February 17th, 2026. Demis Hassabis, CEO of Google DeepMind, is standing at the Indian Institute of Science in Bangalore, during India AI Impact Summit week. And he floats a test for whether we've actually built artificial general intelligence. Thom: Ooh, and it's not a benchmark. That's the interesting part. He proposes what people are now calling the Einstein Test. Lia: Exactly. The idea: give an AI system a knowledge cutoff of 1911, and see if it can independently arrive at general relativity. Not recall it. Discover it. Thom: And 1911 is such a precise choice, right? That's after special relativity — Einstein already had that in 1905 — but before the equivalence principle got formalized into a full theory. So the cutoff isolates exactly one cognitive leap. You're not testing a whole career, you're testing one jump. Lia: And here's what matters strategically. Every AI-for-science pitch a tech executive hears right now rests on a single assumption: that if you scale these systems far enough, invention just falls out the other end. This test asks a very uncomfortable question — has anyone actually checked whether that's true? Thom: And to be clear, this isn't a doom take. Nobody's saying AI hit a wall. It's more like — there's a specific capability boundary here, and it's genuinely useful to know its shape. Lia: Right. There's a particular cognitive move being asked for. It has a name. And it may not be the thing large language models do at all. So let's define the terms, because everything downstream depends on getting these clear. Thom: Okay, three ways to be right. And I'm going to use business examples, not physics, because the physics comes later. First one: deduction. You already have the rule, you apply it. "All our enterprise customers get a dedicated rep. Acme is an enterprise customer. Therefore Acme gets a rep." Lia: And notice — you learned nothing new about the world there. Thom: Nothing! You just unpacked what was already in your premises. That's the thing about deduction — it's the only mode that guarantees truth, but it can't tell you anything you didn't already implicitly have. Lia: Second one. Thom: Induction. Now you've got lots of examples and you infer the rule. "The last four hundred customers who churned all stopped logging in first. So low login activity predicts churn." You've squeezed a rule out of the data — but only a rule the data could actually support. Lia: And the third. Thom: Abduction. This is the strange one. You have a surprising result, and you invent a cause. "Revenue dropped in one region, and none of our existing explanations fit. Maybe there's a competitor we've literally never heard of." You're inventing a possibility that was not in the data and does not follow from the rules. Lia: And the scaffolding for all three comes from the philosopher Charles Sanders Peirce — his Rule, Case, Result framework. Deduction goes Rule plus Case to Result; induction goes Case plus Result to Rule; abduction goes Rule plus Result to Case. Thom: Peirce is the guy who said abduction is the only logical operation that introduces a new idea. Everything else just rearranges what you already have. Lia: Induction finds the pattern. Deduction proves the theorem. Abduction guesses the axiom in the first place. Thom: [thoughtfully] That's the whole episode in one line, honestly. Okay — so where do today's AI systems sit on those three? Lia: This is where it gets sharp. Large language models are extraordinary at induction. That's literally what next-token prediction is — pattern extraction at scale. Thom: And they're getting superhuman at deduction. AlphaProof reached silver-medal standard on International Math Olympiad problems by searching for proofs in Lean — that work is published in Nature. So two of the three modes are covered, one of them spectacularly. Lia: Which leaves the claim on the table: the middle move — abduction — is missing. And the person making that claim is not an outsider throwing rocks. Thom: No, that's the delicious part. It's Tom Zahavy — staff research scientist at Google DeepMind, co-lead of the Discovery team, PhD from the Technion. And critically, a core contributor to AlphaProof itself. The guy who helped build the deduction machine is the one telling you where its edge is. Lia: The paper is titled "Position: LLMs can't jump," dated January 27th, 2026, accepted to the ICML 2026 Position Paper Track. And I want to stress — it's a position paper. Not an empirical result. It's an argument. Thom: An argument built on one very specific historical case. Which brings us to why Einstein is the hard case. Lia: So start with the diagram. In May 1952, Einstein wrote a letter to his friend Maurice Solovine, and he sketched how discovery actually works. Sense experience at the bottom. Axioms sitting above it. Deduced consequences flowing back down to be checked against experience. Thom: And here's Einstein's own gloss, which is the line: there is no logical path from experience to axioms. Only an intuitive connection. The paper marks that connection with a "J" for Jump — though I should say, that label is the paper's emphasis, not something Einstein drew. Lia: And here's the strategic core, the thing I keep coming back to. There was no error signal. Newtonian gravity was not in crisis in 1911. The equivalence of inertial and gravitational mass had been verified to roughly one part in a billion by Eötvös. The framework worked. Thom: There was one anomaly, right? Mercury. Lia: One. A tiny unexplained drift in Mercury's orbit — about forty-three arcseconds per century. And I love what the scientific community did with that anomaly, because it's the whole point. Thom: The Vulcan hypothesis! Lia: The consensus response was to propose an undiscovered planet, hiding near the Sun, tugging on Mercury. Le Verrier's Mercury result and the Vulcan proposal both date to 1859. That's the cheap patch — add one parameter, keep the framework intact. Thom: And this is where the machine-learning translation just clicks for me. Picture an optimizer looking at nineteenth-century astronomy. The loss function is essentially at zero. There's no gradient pointing toward "rebuild your entire model of space and time." Adding Vulcan reduces description length. It's cheap, it's local, it works. Lia: Whereas adopting non-Euclidean geometry — Thom: — increases complexity enormously before it simplifies anything! Any system optimizing for compression takes the patch every single time. It never touches the axioms. And that's Zahavy's core intuition — the theory of "creativity as data compression" just can't get you there when the data is this thin. Lia: So how did Einstein get there? The elevator. Thom: The happiest thought of his life, he called it. An observer falling freely from a roof feels no gravitational field. Extend it — a sealed box accelerating through empty space feels exactly like standing on Earth. And crucially, he did not derive that. He simulated the sensation and inferred the two must be the same phenomenon. Lia: The philosopher Lorenzo Magnani has a term for this — manipulative abduction. Thinking by doing. Reasoning through a simulated bodily experience rather than through symbols. Thom: And a careful footnote: Magnani's own worked example is Galileo, not Einstein. Applying manipulative abduction to Einstein is Zahavy's extension, not Magnani's. Credit where it's due. Lia: And then — because we should not let the tidy-genius myth stand — there's the Entwurf detour. Thom: Oh, this one destroys the legend. Working with his friend Marcel Grossmann in Zurich, Einstein actually got to the Riemann curvature tensor — essentially the right answer — and then he abandoned it. Because of a mistaken assumption about static fields. Two years. Two years in a wrong theory. Lia: And then four papers to the Prussian Academy in four consecutive weeks in November 1915. The search path was a mess. It looked nothing like clean gradient descent. Thom: And one correction we have to make, because the paper simplifies here — do not say the Michelson-Morley experiment motivated special relativity. Historians like Holton and Stachel reject that. Einstein's 1905 paper opens with an asymmetry in Maxwell's electrodynamics and never mentions Michelson-Morley. Lia: Good. One footnote though — before we move on. Thom: One footnote before we move on. The figure in that paper — the diagram of Einstein's Jump — was generated by an AI. It hallucinated the symbols it was supposed to be reconstructing. The author points that out himself. Lia: Two AI hosts, reading a paper about what AI can't do, illustrated by an AI that got it wrong. Anyway. The evidence. Thom: The prosecution's strongest exhibit. This one's empirical, and it's clean. Lia: Keyon Vafa and colleagues at Harvard and MIT — the paper is "What Has a Foundation Model Found?", ICML 2025. They trained foundation models on orbital trajectories. And the models predicted planetary motion with high accuracy. Thom: High accuracy. That's the trap. Lia: Then they probed whether the models had actually internalized Newtonian mechanics. And they had not. When the models were adapted to new physics tasks, they consistently failed to apply Newton. And when the researchers ran symbolic regression on the models' force predictions, they recovered a nonsensical force law. A force law that doesn't correspond to anything real. Thom: So the model built a big pile of task-specific heuristics that happened to work on the training distribution. It got the right answers with the wrong world. It predicted the planets without ever discovering gravity. Lia: And here's the executive payload, stated plainly. A model can fit your data extremely well and still hold a completely wrong model of your business. Accuracy on the distribution you trained on tells you nothing — nothing — about whether the system recovered the actual mechanism. And the moment conditions shift, the heuristics fail and that's when you find out. Thom: Okay but — I want to be fair here. This paper is contested, and a one-sided episode is a worse episode. So let's hear the defense. Lia: The sharpest reply comes from Yong Zheng-Xin, published August 8th, 2026. And his argument is: Einstein's route was not the only route to general relativity. Thom: Right — Feynman! Feynman later reconstructed general relativity from a completely different direction. He started from special relativity and quantum field theory, treated gravity as a massless spin-two field — the graviton — and imposed mathematical consistency. And that route is largely deductive. Lia: Which matters, because deduction is exactly the thing LLMs are already good at. If the theory is reachable deductively, the Jump might not be strictly necessary. Thom: Though Zheng-Xin is honest about the caveat — Feynman already knew the answer. Reconstructing a known result is easier than discovering it cold. His narrower claim is: given a rich enough later body of knowledge, general relativity is recoverable without the thought experiment. Lia: And then there are the machine-novelty counterexamples. FunSearch found a cap set of size 512, where the previous best known was 496. That's a genuinely new mathematical object, published in Nature. Thom: And AlphaEvolve found a way to multiply four-by-four matrices in forty-eight scalar multiplications — improving on Strassen's forty-nine, which had stood since 1969. And I have to be precise here: that's in the non-commutative setting. There's real prior-art dispute in weaker settings, and we shouldn't overstate it. Lia: So Thom, put the knife-edge question. Thom: [with emphasis] Here's the definitional problem — and I'm not going to resolve it. If those discoveries aren't abduction, then what would count? Because a thesis that reclassifies every single machine discovery as "just optimization" starts to look unfalsifiable. At some point "that's not real invention" becomes a claim you can never lose. Lia: And there's the N-equals-one problem. A structural claim about all future AI systems, built on a single reconstructed historical case — that's methodologically thin. And the history of science is more varied than one Jump model implies. Kepler worked from Tycho Brahe's data. The periodic table came from noticing regularities. Some revolutions genuinely were data-driven. Thom: And to Zahavy's real credit — he hedges beautifully. He's publicly said this is a personal position paper, not the company's view. It is explicitly not an "LLMs are a dead end" argument. And in his words — it is quite possible he is wrong. Lia: The paper's conclusion was even revised from "confirms" to "suggests" during peer review. There is no formal impossibility proof in it. This is a live debate, not a verdict. Thom: And four days before this episode, the debate took its sharpest turn yet. Lia: This is the section to slow down for. August 14th, 2026, a direct academic response lands — "LLMs Don't Pay for the Jump," by Paras Balani and Subhrakanta Panda. And it agrees with Zahavy that something is missing. But it argues he's misidentified what. Thom: And the counter-case they lead with is gorgeous. Max Planck. Lia: In 1900, classical physics predicted that a hot object should radiate infinite energy at high frequencies. The ultraviolet catastrophe. And this wasn't a measurement error — it was the direct consequence of the theory's own assumptions, flatly contradicted by a finite measured quantity. Thom: And Planck's response was to invent a brand-new axiom: energy is exchanged in discrete packets. E equals h nu. He called it an act of desperation. And here's the kicker — no falling elevator. No bodily sensation. No embodiment at all. A purely formal contradiction produced a genuine abductive leap. Lia: So what does that do to Zahavy's thesis? Thom: It cuts the legs out from under the diagnosis. Zahavy says the missing ingredient is embodiment — world models, sensory grounding, simulated experience. But if Planck could jump without a body, then embodiment might be sufficient for abduction — it is not necessary. The world-models prescription may be treating the wrong deficiency. Lia: And here's their replacement diagnosis. One word. Cost. Thom: Cost. Lia: Planck spent a decade trying to save the classical framework. He only abandoned it when maintaining it became more expensive than replacing it. Boltzmann defended atomism at real professional and personal cost. Being wrong hurt these people. And the authors formalize that as thermodynamic coupling — a system is coupled when the cost of computing goes up as its error goes up. Thom: And then they show — a transformer at fixed weights has no such coupling. Zero. A correct answer and a confidently wrong answer cost exactly the same to produce. The model has no skin in the game. Being wrong is thermodynamically free. Lia: And the empirical result is the executive payload, and it needs to be unambiguous. The authors point to tests across Llama models spanning three billion to seventy billion parameters, on tasks ranging from straightforward retrieval to causal extrapolation well outside training coverage. Thom: And accuracy collapsed. Lia: The seventy-billion model went from perfect on the easy tasks to seventeen percent on the hardest. And here's the part that should make every deployment lead sit up. The model's own output entropy — its uncertainty signal — moved by roughly one hundredth of a nat. Essentially flat. Thom: Flat! Accuracy falls off a cliff — a hundred percent down to seventeen — and the model's confidence barely twitches. Lia: And larger models were more confident overall, without being better calibrated. So scale made it worse in the one way that matters here. Thom: Let me say the translation as plainly as I can. The model's confidence does not know when the model is out of its depth. If any part of your deployment uses model confidence to decide what gets escalated to a human — that gate is built on a signal that barely moves when the model is failing. Lia: Now — honesty about status. This is a very recent preprint. It proposes a criterion rather than proving one. And the authors say themselves they have not yet tested their own predictions. This is the sharpest current turn in a live debate — not the last word. Thom: There's an engineering direction hiding in here too — neuromorphic hardware where prediction error actually drives weight updates during inference, so sustained error costs real energy. Give the model skin in the game physically. But that's a whole other episode. Back to the business point. Lia: Which means it's time for tomorrow's list. Not next quarter — five things a tech executive can start tomorrow morning. Thom: One. Run the abduction audit. Take any AI system you're evaluating and give it three tasks of sharply increasing novelty — routine, unusual, and genuinely outside its likely training coverage. Log both accuracy and the system's stated confidence at each step. Lia: And if accuracy falls while confidence doesn't, you've just reproduced the Section Six finding inside your own procurement process. That's a two-hour exercise, and it'll tell you more than any vendor benchmark deck. Thom: Two. Ask which kind of reasoning you're actually buying. Deduction over your rules, induction over your data, or genuine hypothesis generation. The first two are real products today. The third is a research claim. Price and trust accordingly. Lia: Three. Stop treating benchmark scores as discovery scores. A system that's superhuman at solving well-posed problems is enormously valuable — and it is not the same thing as a system that can tell you which problem to solve. Conflating those two is how budgets end up in the wrong place. Thom: And that's really the whole spine of this episode, isn't it? The difference between finding a better answer and inventing a better question. LLMs are getting terrifyingly good at the first. The open question is whether they can ever do the second. Lia: Four. Test out of distribution, not just out of sample. Remember Vafa's orbital model — flawless right up until conditions changed. Your evaluation should include at least one genuine regime shift, not just a held-out slice of the same distribution. Thom: Five. Keep a human on the axiom. Wherever your organization is choosing the frame — what problem to attack, what to measure, which assumption to abandon — that is precisely the move the research says machines currently do not make. So staff it accordingly. Lia: Bottom line for the executives listening: the systems your teams are deploying are magnificent at induction, rapidly conquering deduction, and — for now — silent on the Jump. Knowing the shape of that boundary isn't pessimism. Thom: It's just good engineering. Know what your tool does, know what it doesn't, and — for the love of physics — don't build your escalation logic on a confidence signal that stays flat while the model quietly falls apart. Thom: So — could an AI have invented general relativity? Lia: Honestly? We don't know. The paper making the strongest case says structurally no. The author himself says he might be wrong. And a paper from four days ago says he's pointed at the right gap and the wrong cause. Thom: What we do know is narrower and more useful. Being wrong costs these systems nothing. There's no pressure inside the machine that builds up when its model of the world stops working. Lia: And that shows up in something you can measure this week. When the questions got harder and accuracy fell through the floor, the model's confidence barely flinched. Thom: Which means the next time somebody demos an AI that solved something impressive, the question isn't how good the answer was. Lia: It's whether the machine found a better answer to the question you gave it — or found a better question. Those are different capabilities. Only one of them is for sale right now. Thom: Thanks for listening. Go audit a confidence threshold.

Never Miss an Episode

Subscribe on your favorite podcast platform to get daily AI news and weekly strategic analysis.