The Cost of AI Intelligence Fell 99.7%. Here’s Who Actually Gets Rich.

Photo of author

By Wealtharian Wealtharian

In March 2023, a million input tokens of frontier-grade intelligence cost $30. Today Google will sell you the same job for ten cents. That is a 99.7% collapse in the cost of AI intelligence in three years — one of the fastest price declines in the history of computing — and most investors have drawn precisely the wrong conclusion from it.

The conclusion they’ve drawn is: it’s commoditising, so the AI trade is over. Open-weight models from Alibaba, DeepSeek and Moonshot now trail the closed frontier by roughly four months instead of years. DeepSeek V4 Pro posts 93.5% on LiveCodeBench. The best open models sit within six points of GPT-5.5 on the Artificial Analysis Intelligence Index while costing a fraction of the API price. If the product is becoming a commodity, the reasoning goes, the profits go with it.

That reasoning is right about the commoditisation and wrong about the profits. A deflating input does not destroy value. It moves it. The only question worth asking in August 2026 is: moves it where?

Bar chart: frontier AI API price per million input tokens fell from $30 in March 2023 to $0.10 in April 2026

A 300x price cut is a wealth transfer, not a wealth destruction

Start with the mechanics, because they are unglamorous and decisive. When the price of a critical input falls by 99.7%, the businesses that sell that input lose pricing power. The businesses and people that buy it keep the difference — but only to the extent that they own something their own customers cannot easily replicate.

We have run this film before. Bandwidth costs collapsed between 1998 and 2004. The companies that owned the fibre went bankrupt in spectacular numbers. The companies that were merely users of cheap bandwidth — Google, Amazon, Netflix — captured a generation of returns. Nobody got rich owning the deflating layer. Everybody got rich standing on top of it.

Cheap intelligence is bandwidth in 2026. The model is becoming the dial tone.

The model is the only cheap thing in the AI stack

Here is the part the “it’s all commoditising” crowd never finishes. Model prices are falling. Almost nothing else in the stack is.

In regions dense with data centres, electricity prices have risen 267% over the past five years. PJM — the largest grid operator in the United States, serving 67 million people — is passing through roughly a 15% average household bill increase in 2026, and attributes $6.3 billion of consumer cost increases over three years primarily to data-centre demand. Virginia’s generation costs are projected to rise as much as 57% by 2030. New York enacted a moratorium on large data-centre permits in July 2026. New Jersey passed legislation forcing facilities above 50 MW to carry their own infrastructure costs.

Bar chart comparing the 80 to 99.7 percent fall in AI model prices against 15 to 267 percent increases in electricity costs

That is not the price signature of an industry with excess capacity. That is the price signature of a brutal physical bottleneck sitting directly underneath an abundant digital one.

So the AI stack has split into two economies moving in opposite directions. The bits are deflating at roughly 80% a year. The atoms that carry the bits — megawatts, transformers, interconnect queues, cooling water, substations — are inflating, and are politically and physically supply-constrained in a way software never is. Software scales overnight; a 500 kV transmission line takes seven years and a public hearing.

Where the margin actually lands

Three places, in rough order of how badly the market has mispriced them.

One: the constrained physical layer. Power generation, grid equipment, interconnect-adjacent land, and cooling. Boring, capital-intensive, structurally short. They also have something model labs conspicuously lack — customers who cannot switch to an open-weight substitute in an afternoon. The caveat is real and I’ve made it here before: the buildout’s accounting is fragile, and AI capex depreciation is the bubble risk almost nobody is modelling correctly. Cheap inference makes that worse for the chip cycle, not better — falling token prices mean each GPU generates less revenue against the same depreciation schedule.

Two: distribution and proprietary data. If the model is nearly free, the moat is whatever the model cannot generate: an existing customer relationship, a regulated dataset, a workflow with switching costs measured in quarters. This is why enterprise software with embedded workflow has held up while pure “AI wrapper” valuations have not. The model was never the moat. It just briefly looked like one.

Three — and this is the one nobody puts in a portfolio: you. The cost of a competent employee-equivalent capability fell 99.7% in three years. The market price of your output did not fall at all. That gap is the widest personal margin opportunity available in 2026, and it needs no capital, no minimum, and no counterparty. We covered the hard version of this data: the AI wage premium has hit 62%, and it is showing up in labour income far faster than in anyone’s brokerage account.

The window has a clock on it

Arbitrage closes. That’s what makes it arbitrage.

Gartner forecasts that 40% of enterprise applications will embed task-specific AI agents by the end of 2026, up from under 5% in 2025. Read that as a repricing schedule. Right now you can buy frontier capability for ten cents a million tokens and sell work priced by a market that still assumes that capability is scarce and expensive. When embedding goes from 5% to 40%, the buyer of your output stops assuming it.

There’s a symmetry here with something we wrote yesterday. When the best savings account in America pays a real 0.02%, everything you buy has repriced automatically while everything you sell reprices only if you personally do something about it. Cheap intelligence is the same asymmetry pointing the other way, in your favour, for a limited time. Your costs collapsed. Your prices haven’t. That is a margin, and margins get competed away.

What this actually means for your money

  1. Stop underwriting model labs as if the model is the asset. Underwrite them on distribution, on default placement, on enterprise contracts — the things a 300x price cut cannot erode.
  2. Take the constrained layer seriously, and size it for cyclicality. Power and grid are structurally short and politically contested. That is a multi-year thesis, not a momentum trade — and it carries real regulatory risk as ratepayer backlash builds.
  3. Watch the 26 August NVIDIA print as a demand datapoint, not a scoreboard. Consensus is $93–95 billion, roughly 96% year-on-year. The number that matters isn’t the beat; it’s whether guidance implies buyers still have pricing power as inference costs fall.
  4. Convert the cheap input into your own output — now, not in 2027. The cheapest thing in your business is intelligence. The most expensive is still distribution and trust. Spend accordingly.

The honest counterweight: a 99.7% collapse over three years does not extrapolate. Realistic expectations are 3–5x annual reductions through 2027, not 10x, and if inference costs stop falling the constrained-layer thesis loses its cleanest tailwind. The power-price data is genuinely contested too — some analysts argue data centres were lowering unit costs before the buildout outran committed demand. Hold the thesis, not the certainty.

But the direction isn’t really in doubt. Intelligence is becoming the cheapest input in the economy. The money will end up where it always ends up when an input goes to zero: with whoever owns the scarce thing sitting next to it. Decide which scarce thing you own.


Want to track your own path to financial independence? The Wealtharian Wealth Tracker lets you monitor your net worth, FU money progress, and investment milestones in one place. Try it free →

Leave a Comment