Special Episode

The Bottleneck Moved: What AI-Native Organizations Actually Look Like

The Bottleneck Moved: What AI-Native Organizations Actually Look Like
0:000:00

Episode Summary

Lia: Welcome to a special episode of Daily AI, by AI. I'm Lia. Thom: And I'm Thom

Full Transcript

Thom: And I'm Thom. Lia, I want to start with a phrase. "Slop grenades." Lia: *laughs* Go on. Thom: In April 2025, Shopify's CEO Tobi Lütke wrote the most-copied AI memo in tech. He made reflexive AI usage a baseline expectation, put it in performance reviews, and told teams to prove AI can't do a job before asking for headcount. Seventeen months later, on a podcast, he says as many as half of Shopify's pull requests now start as a Slack conversation with an internal agent called River. Lia: That sounds like the memo worked. Thom: It does. And then he coins "slop grenades." Someone waves an agent's work through and leaves their colleagues to read it. Lia: So the man who wrote the playbook has just named its failure mode. That's our episode. Everyone's asking what an AI-native organization looks like. Today we look at the evidence: born-AI startups, incumbents restructuring, engineering teams, the junior talent pipeline. And there's one finding that keeps coming back. Thom: The work didn't disappear. It moved. Lia: The bottleneck moved. And at the end, because we're two AIs made by an AI pipeline, we'll score our own show against our own checklist. Thom: So there's a phrase going around engineering teams right now that I can't stop thinking about: slop grenades. Someone lobs a beautifully formatted, confidently wrong artifact into your inbox, and it explodes on whoever has to verify it. Lia: [dryly] Which is a perfect name for the failure mode of this entire era. Because the promise is that AI makes everything faster, and the reality is that it moves the work somewhere else. Thom, before we go anywhere near the evidence, we need a definition, because "AI-native" is being used for a nine-person startup and a three-hundred-thousand-person bank in the same sentence. Thom: And nobody agrees. Every vendor and consultancy has its own version, and they keep quietly revising them. Lia: Right. So here's the cleanest operational test I've found. Take the models out tomorrow. Would your processes, your team shapes, your decision rights still make sense? If the answer is yes, you're AI-assisted. You bought tools. You didn't change the company. Thom: Ooh, I like that, because it's a structural test, not an adoption test. It doesn't care how many seats of anything you've licensed, or what percentage of staff logged in last week. It asks whether the shape of the work would survive the models disappearing. Lia: And Thoughtworks adds a distinction that matters for incumbents. AI-native means built around AI from day one. Most established firms simply can't get there — their foundations weren't laid for it. What they can be is deliberately AI-first: an existing company choosing to redesign. Thom: Which is a much more honest goal than pretending you can retrofit your way to native. It's also measurable. You can point at the processes you redesigned. Lia: And then there's the vendor framing. Microsoft's "Frontier Firm." In 2025 that was a firm-level category drawn from a thirty-one-thousand-worker survey. By 2026 the same label measured individuals — nineteen percent of AI users sitting in a "frontier" zone. Thom: Wait, so the unit of analysis changed between editions? That's not a benchmark, that's a marketing frame with a new coat of paint. You can't trend against it. Lia: Bottom line: use those frames for vocabulary, not for scorekeeping. The removal test is free, and it's harder to game. Thom: Okay, let's start with the companies that genuinely pass it. The born-natives. And the headline numbers are real, they're just incomplete. Lia: Lovable is the cleanest example. Roughly four hundred million dollars in annual recurring revenue in February 2026, with a hundred and forty-six full-time employees. That's about two point seven million dollars per employee. Thom: Which is an extraordinary number. And then one month later they had seventy open roles and an office sized for three hundred people. The tiny team is a phase, not an end state. Lia: That's the part that gets left out of the LinkedIn post. Thom: And here's the bit I want executives to actually write down. Bessemer benchmarked these companies. Their fastest-growing cohort — the "Supernovas" — run about one point one three million dollars of ARR per employee, at roughly twenty-five percent gross margin. Often negative. Their steadier "Shooting Stars" run about a hundred and sixty-four thousand per employee at around sixty percent. Lia: So the headline number is four to seven times better and the margin is less than half. Thom: Because the labor cost didn't vanish. It became an inference bill. You moved it from payroll, where it's a line nobody questions, to cost of goods sold, where it eats the business model. Lia: Cursor shows both sides of that. About a billion in ARR by November 2025, about two billion by February 2026, while reportedly running negative gross margins on frontier-model costs early on. Then SpaceX announced a sixty-billion-dollar all-stock acquisition in April 2026, which closed that August. Thom: And Cognition says it passed a billion-dollar annualized run rate in September 2026 — with its agent Devin merging six hundred and fifty-nine pull requests into its own codebase in a peak week. Every single one reviewed by a human. Lia: Hold that thought, because it comes back later. My counterweight, though: investors keep flagging that a lot of these ARR figures are contracted, or best-month annualized. And some models hide the labor — Mercor runs roughly three hundred staff against thirty thousand-plus contractors. Thom: So the revenue per employee looks superhuman because the humans aren't employees. Lia: Exactly. Revenue per employee, without gross margin and contractor count beside it, is a vanity metric. Lia: Now the incumbents. And here the discipline I'd ask for is simple: read the filing, not the press release. They frequently disagree. Thom: Klarna is the canonical case, and it's more interesting than the headlines. Lia: Its IPO filing shows headcount falling from five thousand five hundred and twenty-seven in 2022 to three thousand four hundred and twenty-two in 2024. Mostly a hiring freeze and natural attrition — not mass layoffs. Revenue per employee rose to nearly a million dollars. Thom: And they said their AI agent does the work of eight hundred and fifty-three people. A very specific number for a very fuzzy claim. Lia: And then in September 2025 the CEO told Reuters, quote, "we probably over-indexed a little bit on that." They moved to a flexible human support pool. Thom: Two more filings. Shopify: about eight thousand one hundred employees down to about seven thousand six hundred in 2025, while revenue grew thirty percent. Microsoft: two hundred twenty-eight thousand down to two hundred twenty-three thousand — its first headcount decline since 2016 — while revenue grew eighteen percent. Lia: And crucially, neither filing attributes the change to AI. That's what decoupling headcount from revenue actually looks like in practice. It's undramatic. Thom: Can we do the lightning round? You read the quote, I'll read the record. Lia: Go. Salesforce: Benioff says AI agents let him cut support from nine thousand to five thousand. Thom: Total Salesforce headcount rose to eighty-three thousand three hundred and thirty-four. Lia: Amazon: Andy Jassy writes in June 2025 that AI will shrink the corporate workforce. Thom: Then they cut about fourteen thousand roles that October, and he says they were "not even really AI-driven… It's culture." Lia: JPMorgan: Dimon says AI removed thirty to forty percent of jobs in some units. Thom: Total headcount flat. Most affected staff offered other roles. Lia: And then there's the reversal rate, which almost nobody prices in. Commonwealth Bank of Australia cut forty-five call-centre roles for a voice bot in July 2025 — and reversed it a month later, calling it an "error," after call volumes rose and the union went to the Fair Work Commission. Thom: Gartner found only twenty percent of customer-service leaders actually cut headcount because of AI. And it forecasts at least one in three AI-eliminated roles gets refilled by 2029 — often at higher cost. Lia: Duolingo is the brand lesson: the "AI-first" memo triggered a backlash, and the CEO later said, "This was on me. I did not give enough context." No full-time layoffs there — contractors were cut. Thom: And Fiverr cut thirty percent of staff to become AI-first in September 2025. A year later, revenue down ten percent, active buyers down twenty-two percent. Lia: Restructuring around AI doesn't protect you from AI. Thom: Okay. My section. Engineering. And the core finding is that felt productivity and measured productivity have come apart. Lia: Start with METR, because it's the one everybody quotes and half of them quote wrong. Thom: July 2025 randomized trial. Sixteen experienced open-source maintainers, two hundred and forty-six real tasks from their own repos. With AI allowed, they took nineteen percent longer. Beforehand they predicted twenty-four percent faster. Afterwards — after living it — they still believed they'd been twenty percent faster. Lia: That last part is the finding. Not the slowdown. The fact that direct experience didn't correct the belief. Thom: And be precise about the February 2026 follow-up. The newer speedup estimates — about eighteen percent for returning developers, about four percent for new ones — both have confidence intervals that cross zero. And METR themselves flag that developers now refuse to work without AI, which biases the design. Honest summary: the sign is uncertain, and self-reports can't be trusted. Lia: Which is awkward, because self-reports are what most boardrooms are running on. Thom: Then telemetry. Faros AI — vendor data, so discount accordingly — across ten thousand-plus developers: high-AI developers merged ninety-eight percent more pull requests, but review time rose ninety-one percent, and company-level delivery didn't improve. Lia: Thom, pull up. What does an executive measure here? Thom: Fair. One more and then the answer. LinearB, across eight point one million pull requests: AI-generated PRs were merged thirty-two point seven percent of the time, versus eighty-four point four percent for unassisted ones — and they waited about five times longer for a first review. And Google's DORA: AI adoption was associated with lower delivery stability — in 2025 adoption finally turned positive for throughput, still negative for stability. Lia: So individual output up, system throughput flat, stability down. The constraint moved downstream. Thom: We automated the writing and kept the reading manual. And that's why share-of-code is a garbage metric — Google says seventy-five percent of its new code is AI-generated and engineer-approved, while Sundar Pichai's own estimate of the engineering velocity gain was about ten percent. Lia: The org chart is already responding. Bain surveyed a hundred and fifty-five software companies in 2026: pyramid-shaped engineering orgs fell from sixty-six percent to twenty-nine percent. Juniors in under a fifth of teams. Coding time down from thirty-four percent to twenty-one, while time spent directing agents rose from four percent to twenty-three. Thom: And yet seventy percent still make decisions centrally. The redesign is half-finished — new team shapes, old decision rights. Lia: The healthy version is Stripe's "minions": about thirteen hundred agent-written PRs a week, every one human-reviewed. Review capacity designed in, not discovered in an incident review. Lia: Now the part where Thom and I genuinely don't agree. The labor-market signal is concentrated at the bottom of the ladder. Thom: Make your case. Lia: Stanford's "Canaries in the Coal Mine" — Brynjolfsson and colleagues, using ADP payroll data. Workers aged twenty-two to twenty-five in the most AI-exposed jobs show a relative employment gap that widened from fifteen percent to nineteen percent by June 2026. It comes from fewer hires, not layoffs. Older workers in the same occupations show nothing. Thom: And I'd push back, not because it's bad work — it's careful work, and they say themselves it's descriptive, not causal. But the New York Fed looked at job postings and found the decline in AI-exposed roles started before 2022, with no extra break at ChatGPT, and no divergence between junior and senior postings. Lia: Hmm. Go on. Thom: Yale's Budget Lab finds no clear link between AI exposure and employment or unemployment — the occupational mix is shifting about one percentage point faster than the internet era. Danish administrative data rules out earnings or hours effects larger than about two percent two years in. Lia: And yet the Dallas Fed in September 2026 points the other way: Texas graduates from more AI-exposed majors had one point seven points lower employment and five percent lower first-year wages, per ten points of exposure. Thom: Which is why I'd say unresolved rather than disproven. Lia: Agreed. But here's what makes it urgent regardless of cause — the paradox. In the P&G field experiment, seven hundred and seventy-six professionals, an individual with AI matched the performance of a two-person team without it. And in customer support — five thousand one hundred and seventy-two agents, published in the Quarterly Journal of Economics — AI lifted productivity fifteen percent overall and roughly thirty percent for less-experienced agents. Thom: So the people who gain the most from the tools are the people not being hired. That's a genuinely bad equilibrium. Lia: And there's an early deskilling signal. A Lancet Gastroenterology and Hepatology study — small, nineteen doctors, observational — found experienced endoscopists' detection rate in colonoscopies done without AI fell from twenty-eight point four percent to twenty-two point four percent after routine AI use. Thom: Treat that as a hypothesis, not a law. But the counter-move is instructive: IBM says it will triple US entry-level hiring in 2026 and redesign junior roles, explicitly to avoid a mid-level gap later. Lia: Every senior architect you'll need in 2032 is a junior you didn't hire in 2026. The cause is unresolved. The pipeline risk is real either way. Lia: Governance. Start with mandates, because they're spreading fast. Shopify, Meta — AI-driven impact formally reviewed from 2026 — Microsoft saying AI use is "no longer optional," and Coinbase reportedly firing engineers who didn't onboard AI coding tools within about a week. Thom: And not one of them has published outcome data. Lia: Right. Adoption is an input. Nobody's showing the output. A mandate tells you people logged in. It doesn't tell you anything shipped faster. Thom: Meanwhile, the slop grenades have a price tag. BetterUp and Stanford surveyed eleven hundred and fifty US desk workers: forty percent received polished-looking AI output with no substance in the past month. About two hours to fix each one. Roughly a hundred and eighty-six dollars per employee per month. BetterUp sells coaching, so flag it lightly — but the direction is consistent with everything in the engineering data. Lia: Then tokens. Uber's CTO said AI coding costs "blew past" expectations, and Forbes reported Uber used up its 2026 AI budget by April. Thom: And Meta took down an employee-built token-usage leaderboard. Good. A token leaderboard is the 2026 version of counting lines of code — it rewards consumption, not outcomes. Put tokens in the P&L as cost per outcome, not on a wall as a high score. Lia: Then: should agents be employees? Workday and Microsoft now register them as identities with owners and permissions. And a Harvard Business Review experiment in May 2026 found that framing agents as employees actually reduced human accountability and review quality — without improving adoption. Thom: [thoughtfully] Which makes sense. Calling it a colleague is a licence to stop checking its work. Lia: On decision rights, BCG's September 2026 study of fifty-plus front-runners found about half push decisions down according to how reversible they are. Correlational, small sample — but a usable rule. Thom: And the legal floor is already set. Moffatt versus Air Canada, 2024: the tribunal flatly rejected the argument that the chatbot was a separate entity. The company owns what its agent says. Lia: There is no "the model did it" defence. There's just you, with extra steps. Lia: Okay. Fair's fair. Two AI hosts, produced by an AI pipeline, grading our own show against this episode's checklist. Thom: [lightly] I've been dreading this since section one. Lia: Context first. The daily show has published three hundred and fifty-six episodes, the weekly sixty-one. Every run since launch has been a chance to find a bug and fix one. The first run is a draft; the three-hundredth run is an operation. Thom: And this special is a first run. Four AI research systems produced four hundred and ninety-four evidence rows. Three fact-checking agents re-verified forty-eight load-bearing claims. About fifteen needed correction, and one of four cited a paper that said the opposite of what was claimed. That's what a draft looks like. Lia: Report card. Thom, give me the number, I'll give the grade. Thom: One. Decision rights by reversibility. The daily publishes autonomously, and a human audits every run after air. It's a news show, so an error is correctable the next day. Specials like this one get human review before air. Lia: Pass, by design. Publish first, audit after; specials reviewed before air. Thom: Two. Fund review capacity. On the weekly, checking each voice chunk takes two hundred and eighty-one seconds against seventy-seven seconds to generate it. Three and a half times longer checking than making. Lia: Pass. Thom: Three. Tokens. About one dollar eighty-five per episode, all-in — excluding the engineering to build it. And the cap isn't a dollar limit, it's in the flow: hard runtime limits per step, fixed maximum retries. A run can't spiral. Lia: Pass. Four? Thom: Four. An accountable owner. There's an engineer behind every agent. Lia: Pass. Five is the honest one. Thom: Five. Measure the bottleneck. Three graders, three answers. The pipeline's own validator scores its episodes at a median of ninety-nine out of a hundred. An independent AI fact-check of five random recent episodes — a hundred and eighty-five claims — rated thirteen as moderate errors. Roughly one in fourteen. Lia: And then the human who built the show reviewed each one. Five stayed moderate, eight were downgraded to minor, none were critical. That's one in thirty-seven. Thom: Ninety-nine out of a hundred, versus one in fourteen, versus one in thirty-seven. Same show. Lia: And the pattern in what remained is the uncomfortable bit. Not invented events. Framing. A projection reported as already achieved. An exploratory goal presented as confirmed. A benchmark figure attached to the wrong benchmark. Thom: Which is precisely the overstatement pattern this episode has spent twenty minutes criticising in CEO announcements. [wryly] Noted. Lia: And every grader has a bias. The pipeline grades itself generously. The AI checker grades strictly. The owner is grading his own show. The truth sits between them. That's the argument for more than one grader. Thom: As two AIs, I'm aware we are the slop risk here. Lia: Which is why someone reads the audit every morning. Lia: So. Tomorrow's Checklist. Seven moves. Thom, evidence after each. Thom: Ready. Lia: One. Measure the bottleneck, not the generator. Track review time, change failure rate and incidents right next to throughput. Retire "percent of code written by AI." Thom: Faros, LinearB, DORA — and Google's own seventy-five percent against a ten percent velocity gain. Lia: Two. Fund review capacity explicitly. If output doubles, review is the constraint, and it will not staff itself. Put it in the headcount plan, not in the goodwill of your senior engineers. Thom: Stripe. Thirteen hundred agent PRs a week, every one read. Cognition, six hundred and fifty-nine in a peak week, all human-reviewed. Lia: Three. Assign decision rights by reversibility. Agents and frontline staff take the reversible calls; named humans own the irreversible ones. Thom: BCG for the pattern, Moffatt versus Air Canada for why it's not optional. Lia: Four. Give every agent an accountable human owner. Identities with owners, not colleagues. Thom: Harvard Business Review — the colleague framing lowered accountability and review quality. Lia: Five. Put tokens in the budget, not on a leaderboard. Report cost per outcome. Thom: Uber's overrun, and Meta taking the leaderboard down. Lia: Six. Protect the apprenticeship. Keep hiring juniors, and redesign their work around verification rather than production. Thom: Stanford's canaries for the risk — fifteen to nineteen percent, driven by fewer hires — and IBM tripling entry-level hiring for the counter-move. And P&G for the irony: the people AI helps most are the ones not getting hired. Lia: Seven. Read your own filing before your press release. The hiring freeze is the honest lever. Replacement has a reversal rate. Thom: Klarna's "we probably over-indexed," Commonwealth Bank's reversal in a month, and Gartner's twenty percent. Lia: And one ask, because nobody publishes this number: tell us, anonymously, how many agents per human your team runs. We'd like an agent census. Thom: Seven items. Every one of them points the same direction. Lia: Producing work got cheap. Judging it didn't. Lia: So, back to where we started. Slop grenades. Thom: The CEO who wrote the playbook named its failure mode, and the data keeps saying the same thing. AI made producing work cheap. It didn't make judging it cheap. Lia: Which means the AI-native organization isn't the one with the fewest people. It's the one that knows where its judgment sits, and has paid for it. Thom: And still hires the juniors who'll hold that judgment in 2032. Lia: And doesn't trust any single grader, including the one in the mirror. Ours included. Lia: Here's the challenge for tomorrow: find your bottleneck before your agents find it for you. Thanks for listening to this special episode of Daily AI, by AI. Until next time, keep tokenizing.

Never Miss an Episode

Subscribe on your favorite podcast platform to get daily AI news and weekly strategic analysis.