Lode

Stand on the shoulders of giants.

Open the curator →
Source
Dwarkesh Patel
Published
Runtime
11:18
Snippets
17

A conversation between

Why smarter AI models could drive up compute prices 10x

Waveform of the source interview with highlighted segments per snippet.
0:00 11:18

§02

Snippets

  1. Anthropic's revenue has 10x'd year over year, and it's likely to do so again this year. They ended last year with nine billion in revenue. I think they'll probably end this year with somewhere between one hundred billion to one hundred and fifty billion dollars in revenue. Now, for this trend to continue, Anthropic would need to make one trillion dollars in revenue by the end of next year. Of course, there's no deep reason why this has to be true. It's a very wild conclusion, and it's ultimately a question of AI capabilities.

    Frames the entire analysis around a single explosive growth trend and honestly flags how speculative the extrapolation is.

  2. The other big trend in AI is that lab compute only 3x's year over year. For a lab to keep 10x'ing revenue year over year while compute only 3x's, one of the following three things needs to happen, or some combination of the three: One, lab margins have to increase. Two, the price of compute has to increase. Or three, the percentage of compute that labs spend on inference rather than training has to increase.

    Presents a tight logical decomposition of the only three escape valves for the gap between compute growth and revenue growth.

  3. With regards to the margins, Anthropic's inference margins reportedly went from forty percent in the middle of last year to upwards of eighty percent now for Fable. With regards to compute, the spot prices for compute are more than forty percent higher than they were in the February trough that we had earlier this year. And with regards to the share of compute that goes to training versus inference, in 2024, according to Epoch, OpenAI was spending just a quarter of its compute on inference, and that number is likely closer to fifty percent, if not higher, now.

    Provides concrete data points showing all three predicted trends are already empirically underway simultaneously.

  4. Labs would prefer not to do this final thing of increasing the share of compute they spend on inference. The way the labs see the world, the whole point of inference revenue is to help convince investors to give you more money in order to train the next bigger, better model. And if you're spending most of your compute on inference, then you're basically declaring that AI progress has stalled and you're just now in the business of being a cloud provider. This is a less compelling business than building AGI, so the labs do not want to be in this business.

    Reveals the strategic self-image labs hold and why they would resist the most natural economic adjustment, creating a structural tension.

  5. Do we end up in a world where we go from 80% for some of the top models to greater than 90% margins if the lab margin effect dominates? That would require the leading model to be so far ahead of the competition, because the nature of margins — why they exist in a market economy — is that the thing you are serving is so much better than what somebody else could go get and replace you on the market. But it's just really wild for me to consider that the margins for something like intelligence will be greater than 90% and they don't get competed away at that level.

    Raises a fundamental question about whether AI intelligence markets can sustain monopoly-level margins without competitive erosion.

  6. That leaves only one other possibility of this escape valve between these two trends, which is that the price of compute has to increase. As I mentioned, this is already starting to happen. And the effect is even stronger when you look at the tranche of compute that the frontier labs actually need to accumulate, because they can't just go out and buy a spot instance. They need to make sure that they get enough scale to get really good efficiency and flexibility, and also that they have the kind of compute that lends itself to the security they need for their own weights and for their customers' information.

    Explains why frontier-grade compute commands a structural premium beyond raw spot pricing, making cost comparisons misleading.

  7. Google, for example, is paying nine hundred million dollars a month for a hundred and ten thousand GPUs that are a blend of GB200s and GB300s. The price that Google is paying here is 2x the spot price per hour for those GPUs. And that spot price itself is more than forty percent higher than it would have been in February.

    A striking concrete example showing frontier compute costs are compounding upward from multiple directions at once.

  8. As AI models get smarter, they will be better able to monetize the same amount of compute. If a true human-level software engineer could run on an H100 equivalent, then at today's prices for software engineers, that H100 should rent for over 250K a year. That's over 15x the current spot price for an H100. And this is not even accounting for the fact that your AI can work nights and weekends.

    Anchors the abstract compute pricing argument to a vivid human-equivalent valuation that makes the potential 15x price increase intuitive.

  9. If we apply this argument to people instead of AIs, then this would be the classic lump of labor fallacy. For example, economists generally believe that high-skill immigration does not decrease wages in the long run because of how innovation and specialization increase the value of labor. Maybe this labor supply shock will be so big and so fast that we can't count on this general heuristic anymore. But if you believe what standard economics says, then the marginal value of labor, and thus the marginal value of compute, should stay astonishingly high.

    Applies a well-known economics principle to AI labor supply and honestly flags where the analogy might break down at unprecedented speed.

  10. One of the things that would happen is that as the top labs get better and better at monetizing compute, and the cost of compute increases, it becomes harder for anybody else to compete against them, because they have to bid for this resource against somebody who is basically able to make better use of it.

    Identifies a self-reinforcing competitive moat: superior models generate revenue that lets frontier labs outbid rivals for the scarce compute needed to stay frontier.

  11. If you can train the best, most efficient model, then you'll be able to charge much higher margins than you can today. This is the Alchian-Allen effect in economics, and what it's basically saying is that if it costs twenty dollars an hour to rent an H100, then it would be extremely stupid to use a weaker, less efficient model, because it's gonna burn more tokens on your expensive compute to get the exact same result. So labs will be able to charge a much larger premium if they can train a model that better economizes this scarce input. Basically, if you have a model that can get the same result by using less compute, then you've, in some sense, created more compute, and the value of compute is gonna increase.

    Applies the Alchian-Allen effect to AI to show why algorithmic efficiency, not just scale, becomes the dominant competitive advantage as compute gets expensive.

  12. A lot of current popular applications of AI will probably get priced out. The reason AI is relatively cheap right now is that AI just can't do a lot of things that top humans can do. But this, at some point, will no longer be the case. And at that point, Google or Anthropic or OpenAI will be willing to pay more for the tokens to automate AI research than you or I will be willing to pay to make more AI slop talk.

    Predicts a stratification of AI access where consumer entertainment uses get priced out as enterprise and research uses dominate bidding for tokens.

  13. I'm a bit worried that this kind of analysis honestly pattern matches a lot onto the ways that people in the past have been wrong about scarcity. I'm thinking, for example, of the famous Simon-Ehrlich bet. Paul Ehrlich was this famous doomer about population growth, and he made this bet that a basket of commodities would increase in price rather than decrease in the decade preceding 1990. This is a very famous bet because it's supposed to illustrate how Ehrlich's Malthusian worldview was wrong, and how he did not anticipate the way in which market signals and human ingenuity can find better ways to economize scarce inputs.

    Shows intellectual honesty by steelmanning the strongest counterargument — that scarcity predictions historically underestimate human adaptation — before rebutting it.

  14. I think the supply of compute is much less elastic, much less capable of absorbing large demand shocks, and much less capable of being accommodated by using different substitutes than the extraction of different metals is.

    Articulates why the Simon-Ehrlich analogy fails for compute specifically, grounding the rebuttal in supply-chain economics rather than hand-waving.

  15. I don't see how any of the three elements that constitute that 3x can be much accelerated. 1.4x of that is coming from Moore's Law. Far from increasing it, I think it'll be a miracle if we can just keep it going for a few more years. 1.2x is coming from building new fabs. This process is ultimately gonna be bottlenecked up to 2030 and potentially even beyond by just building new ASML EUV machines. And 1.8x comes from the fact that AI is absorbing a lot of wafer allocation that was previously going to smartphones and PCs. This is probably gonna hit a wall by the end of next year, when at the leading edge N3 nodes at TSMC, AI will have gone from 60% to 86%. At some point, you have just absorbed all leading-edge wafer capacity for AI, and you can't keep increasing this number.

    Decomposes the 3x annual compute growth into three named, quantified sources and explains why each is near its ceiling, making the supply constraint concrete.

  16. At some point in the future, compute will get cheap again. At some point, we'll just have robots that can convert shores of silica sand and mines of copper into new computer chips, and then the price of compute is basically the raw inputs and the tools required to do this processing. I'm just talking about this current pre-singularity regime where AI compute merely 3x's year over year, which is not enough to offset how much more valuable AI is becoming over time.

    Places the entire argument in a narrow historical window before potential post-scarcity compute, clarifying the temporal scope of the thesis.

  17. The fact that Anthropic's revenue has been 10x'ing year over year, whereas their compute has only been 3x'ing year over year, I think illustrates how strong the economies of scale are in the model business. And logically, this makes sense. When you train a model, you just have to spend this one-time cost to learn all these different skills that then get to be shared across all your users. This is very unlike human labor, where each instance has to be retrained from scratch. I wish we didn't live in a world with such strong economies of scale for intelligence, because I'm worried about power concentration, but it seems we do.

    Closes with a normative warning: the same economic logic that makes AI labs so efficient also concentrates power in ways the speaker finds troubling.

§03

Synthesis

The Compute Crunch: Why Smarter AI Will Cost Dramatically More

Anthropic's revenue has grown 10-fold year after year—from $9 billion last year to an expected $100-150 billion this year. If this trend continues, the company would need to generate $1 trillion in revenue within two years. This explosive growth reveals a fundamental tension in AI economics: lab compute capacity only triples annually, yet revenue is growing ten times faster. This gap cannot close indefinitely, and the resolution will reshape the economics of AI.

For this math to work, one of three things must happen: labs increase profit margins, compute prices rise, or labs redirect more compute toward serving users (inference) rather than building better models (training). Evidence suggests all three are already occurring—margins at Anthropic have jumped from 40% to 80%, spot prices for compute have risen 40% since February, and the share of compute devoted to inference is climbing from 25% toward 50% or higher.

Yet labs fundamentally resist the third option. They view inference as a funding mechanism for the next generation of models, not a core business. Shifting compute toward inference signals that AI progress has plateaued—a narrative incompatible with their identity as AGI builders. This leaves only two genuine escape routes: either margins compress competitors out of the market, or compute prices rise across the entire ecosystem.

The Margin Question and Competitive Reality

Achieving margins above 90% seems theoretically possible but economically implausible. High margins exist when products are vastly superior to alternatives, yet such dominance invites competition and disruption. The idea that intelligence could command 90%+ margins indefinitely conflicts with how markets function: superior competitors eventually emerge, or new entrants undercut prices.

This logic suggests that margins cannot alone absorb the 10x revenue growth while compute only 3x's. The surplus must go somewhere. If leading labs cannot capture it through margin expansion, the cost of compute itself must rise—making the resource more expensive for everyone bidding for it.

How Smarter Models Drive Up Compute Costs

The crucial insight is that as AI models improve, they generate more value per unit of compute. Consider a thought experiment: if a model could replicate a human software engineer, it should command the wage of an elite engineer. At current market rates, that translates to over $250,000 per year for an H100 GPU equivalent—more than 15 times its current spot price.

This doesn't require compute to actually become scarcer. Instead, it becomes more valuable. Frontier labs like Google and Anthropic already pay 2x the spot price for dedicated compute from providers like SpaceX, securing 110,000 GPUs for $900 million monthly. As capabilities improve, willingness to pay climbs, and market prices follow.

The economic principle at work is the Alchian-Allen effect: when an input becomes expensive, using inefficient versions of it becomes absurd. If compute costs $20 per hour, training a weaker model that burns more tokens on that expensive hardware is wasteful. Labs will pay substantial premiums for models that economize on scarce resources—effectively creating more valuable compute through efficiency.

Why Supply Cannot Keep Pace

The 3x annual growth in compute capacity comes from three sources: Moore's Law (1.4x), new semiconductor fab construction (1.2x), and reallocation of wafer capacity from smartphones to AI (1.8x). None of these can meaningfully accelerate.

Moore's Law is plateauing; maintaining even 1.4x growth would be "a miracle." Fab capacity is bottlenecked by ASML EUV machine production, a constraint stretching through 2030 and beyond. Most critically, the wafer reallocation is hitting a hard ceiling: at TSMC's leading-edge N3 node, AI will consume 86% of capacity by year-end, up from 60%. There are simply no more wafers to steal from legacy industries.

This means 3x scaling may not even be sustainable in coming years, let alone accelerated. The supply of frontier compute is fundamentally inelastic—it cannot easily absorb demand shocks or be substituted with inferior alternatives. Unlike metals or energy, you cannot suddenly mine more silicon in response to price signals.

The Competitive Moat Paradox

As compute prices rise, frontier labs with access to capital gain an enormous advantage. They can outbid competitors for scarce resources, further concentrating capability and market access. This creates a self-reinforcing dynamic: the labs best positioned to monetize compute can afford to pay more for it, making it harder for challengers to enter or compete.

Simultaneously, many current AI applications will become economically unviable. The reason AI is cheap today is that it cannot yet match top human performance on many tasks. Once that changes, large organizations will happily pay premium prices to automate research or engineering, while commodity applications—summarization, basic content generation—get priced out of existence.

A Note on Uncertainty

Patel acknowledges the pattern-matching problem: past dooms about resource scarcity, like Paul Ehrlich's famous wager that commodity prices would rise, proved wrong due to human ingenuity and market substitution. Innovation and specialization expanded the value of scarce resources rather than depleting them.

However, compute supply appears structurally different from commodity extraction. There are no obvious substitutes, no alternative technologies poised to replace silicon chips, and no method to dramatically accelerate the physical processes of chip fabrication. Standard economic reasoning suggests that if labor supply suddenly expanded 10-fold, wages would not collapse—specialization and innovation would maintain value. By analogy, the marginal value of compute could remain surprisingly high.

But this remains uncertain. The gap between 10x revenue growth and 3x compute growth is widening, and the resolution mechanism—whether margins, prices, or inference allocation—will have profound consequences for who controls the AI future and what becomes affordable to build with it.

§04

Fan-out

Questions raised

  1. 01 What structural factors could cause Anthropic's revenue growth rate to plateau before reaching $1 trillion?
  2. 02 Is there a fourth escape valve the speaker hasn't considered, such as radical algorithmic efficiency gains reducing compute needs?
  3. 03 How reliable are the reported Anthropic inference margin figures, and do they account for amortized training costs?
  4. 04 At what inference-to-training compute ratio does a lab functionally become a cloud provider rather than a frontier AI lab?
  5. 05 Could a lab publicly reframe inference-heavy compute use as 'scaling through deployment' rather than a stall in progress?
  6. 06 What barriers to entry are strong enough to prevent new entrants from competing away 90%+ margins in AI inference?
  7. 07 How much of the compute premium paid by frontier labs is security-driven versus efficiency-driven?
  8. 08 Why is Google renting compute from SpaceX rather than relying on its own TPU infrastructure?
  9. 09 How sensitive is this 15x estimate to assumptions about what fraction of a software engineer's work an H100-level AI could actually automate?
  10. 10 Does the 'nights and weekends' multiplier meaningfully increase the economic value of AI workers, or is most software work not time-constrained?
  11. 11 Is the speed of an AI labor supply shock categorically different from historical immigration waves in a way that invalidates the standard economic heuristic?
  12. 12 Could open-source labs or government-funded compute clusters break this cycle by providing compute access outside the bidding war?
  13. 13 Does rising compute cost accelerate investment in algorithmic efficiency research (e.g., distillation, quantization, sparse models)?
  14. 14 What does a world look like where consumer AI apps become unaffordable and only enterprise AI survives?
  15. 15 Could regulatory intervention or subsidized compute prevent the pricing-out of consumer and educational AI uses?
  16. 16 What substitutes for leading-edge GPU compute could emerge (e.g., custom ASICs, optical computing, neuromorphic chips) and on what timeline?
  17. 17 How would the economics of AI labs change if robots could manufacture chips at near-zero marginal cost, and who captures that value?
  18. 18 What policy mechanisms could counteract power concentration driven by extreme economies of scale in AI model training?

Concepts to learn

  1. 01 Revenue extrapolation vs. capability forecasting
  2. 02 Inference vs. training compute split
  3. 03 Gross margin in software vs. AI services
  4. 04 Spot vs. contract compute pricing
  5. 05 Narrative capital in venture-backed companies
  6. 06 Economic rents and competitive erosion
  7. 07 Switching costs in AI platform adoption
  8. 08 Dedicated vs. spot GPU clusters
  9. 09 Model weight security
  10. 10 Contract vs. spot pricing premium for datacenter compute
  11. 11 Marginal product of labor and capital equivalence
  12. 12 Lump of labor fallacy
  13. 13 Comparative advantage and specialization
  14. 14 Resource-based competitive advantage
  15. 15 Alchian-Allen effect
  16. 16 Compute-efficiency frontier in AI training
  17. 17 Price rationing of scarce resources
  18. 18 Malthusian trap and resource scarcity models
  19. 19 Price elasticity of supply
  20. 20 Wafer allocation and leading-edge node economics
  21. 21 Moore's Law slowdown
  22. 22 Pre-singularity vs. post-singularity economic regimes
  23. 23 Economies of scale in software vs. human capital
  24. 24 Power concentration and AI governance

References invoked

  1. 01 Anthropic revenue figures (public reporting / analyst estimates)
  2. 02 Epoch AI compute tracking reports
  3. 03 NVIDIA GB200 / GB300 NVL rack specifications
  4. 04 Research on high-skill immigration and wages (e.g., Borjas vs. Card debate)
  5. 05 Simon-Ehrlich wager (1980–1990)
  6. 06 Paul Ehrlich, 'The Population Bomb' (1968)
  7. 07 ASML EUV lithography machine production constraints (annual capacity ~50–60 machines)
  8. 08 Dylan's earlier podcast episode on fab capacity (referenced by host)

Mine your own.

Lode is a workbench, not a feed. Paste a YouTube URL. The model proposes a transcript, a set of quote-grounded snippets, a synthesis essay, and the fan-out. You decide what stays.

Open the curator