Categories
AI

The Arithmetic of the Sold-Out Warehouse

In the spring of 2026, Nvidia reported a quarter in which it sold $81.6 billion worth of chips, wrote it up at a gross margin of nearly 75 percent, and casually mentioned that cloud GPUs were sold out. Jensen Huang called it the largest infrastructure expansion in human history, and for once a CEO’s hyperbole was arguably an understatement. Revenue was up 85 percent from a year earlier. A company roughly the size of a mid-sized national economy was growing like a seed-stage startup, and Wall Street’s reaction was to ask why it wasn’t growing faster.

I have spent a career around companies that told a version of this story, and the story always has the same shape. Something becomes scarce. Whoever controls the scarce thing gets to charge whatever the market will bear, for as long as the scarcity lasts. The interesting question was never whether Nvidia’s chips were good. Everyone agreed they were good. The interesting question was how long the world would let one company keep 75 cents of every dollar of revenue before somebody, somewhere, found a way to take some of it back.

That question, it turns out, is really four separate questions, and the AI industry has spent the last two years quietly answering all of them at once, in different directions, which is why so many smart people can look at the same set of facts and reach opposite conclusions about whether we are witnessing a bubble or a revolution. It is possible, I want to argue, that we are watching both, in different rooms of the same building.

Start with the money. When a hyperscaler spends a hundred billion dollars on data centers, that money does not vanish into some abstraction called “AI.” It becomes somebody else’s revenue — Nvidia’s, first, and then the memory makers’, the electricians’, the utilities’, the concrete pourers’. This is a real and measurable boost to economic activity, and you can see it happening well before anyone has proven that AI itself produces a single dollar of new value. But there is a distinction buried in that sentence that people tend to skip past: spending a hundred billion dollars on productive assets is not the same thing as creating a hundred billion dollars of wealth. The assets still have to earn their keep. Somebody has to use them for something worth more than they cost.

Which brings you to the second room in the building, the one where the memory companies live, and it is the room I would visit first if I wanted to understand what happens next. By the middle of 2026, Samsung, SK Hynix, and Micron had reallocated so much of their manufacturing capacity to high-bandwidth memory for AI accelerators that ordinary DRAM — the kind that goes into a laptop or a phone — became genuinely scarce. Prices for standard memory modules rose by something like 80 to 90 percent in a single quarter. SK Hynix posted an operating margin north of 70 percent. Micron’s profit rose more than sevenfold year over year. Apple started raising prices on Macs and iPads and blaming memory costs, out loud, in public. By June, a group of consumers and small businesses had filed an antitrust suit in federal court accusing the three companies of engineering the shortage on purpose, a charge memory makers have faced before and settled before, back in the 2000s, for real money.

I don’t know whether that lawsuit has merit. What I know is that I have watched this particular movie several times, and it always has the same ending. Scarcity produces extraordinary margins. Extraordinary margins summon capital. Capital builds capacity. Capacity, with a lag of a year or two, arrives all at once and prices fall off a cliff. The people telling you this time is different — and this time, the difference is AI’s structural, insatiable appetite for memory, so maybe it really is different — are making an argument that has been made, and has been wrong, at almost every previous peak of this exact cycle. Building a new fab takes eighteen to twenty-four months. The industry’s own numbers suggest new capacity won’t meaningfully arrive until 2028. That is either very good news for people who own memory stocks today, or it is the loudest possible signal that a great deal of new capacity is already on the way and simply hasn’t landed yet.

Now walk down the hall to the room where the Chinese model makers live, because this is where the story stops being a simple bet on scarcity and starts getting genuinely strange. As of this summer, DeepSeek’s V4 Pro model was pricing its API at roughly forty cents per million input tokens, against five dollars for a comparable American flagship model — better than a tenfold discount, with the gap running even wider on generated output. Alibaba’s Qwen and Moonshot’s Kimi were sitting in a similar band. Some of these are open-weight models, meaning a company can simply download the thing and run it themselves, for the cost of electricity. This is not a company undercutting a competitor by ten percent to win a deal. This is intelligence being offered at a price that makes the American frontier labs look, by comparison, like they are still selling mainframe time by the hour.

If you take that seriously, it forces an uncomfortable question. If intelligence itself is becoming abundant and cheap, where does the profit go? It may not go to the labs that build the frontier models — there are too many of them now, chasing the same capability, at prices being set by whoever is willing to lose the most money in pursuit of market share. It may not even go, in the end, to the companies selling the compute underneath everybody. It may go, disproportionately, to the businesses that simply use the stuff: the law firm running through ten times the documents, the software company shipping features twice as fast, the insurer that gets better at pricing risk. Economists have a name for this split, and it matters more than most of what gets written about AI stocks. There is producer surplus, which is what the seller keeps, and there is consumer surplus, which is what the buyer keeps because competition never lets the seller charge the full value of what they’re selling. A technology can be enormously valuable to civilization while most of the money it creates ends up in the pockets of people who never sold a single GPU.

Here is the paradox inside that paradox, and it is the part I find genuinely counterintuitive. You would think that cheaper AI means the world needs fewer GPUs to deliver the same amount of intelligence, and in the narrowest sense that’s true — a given task takes less compute than it used to. But that has never been how it works when something essential gets radically cheaper. Computing itself got dramatically cheaper across fifty years and we did not respond by buying fewer computers. We put computers in everything, including things that had no obvious business containing a computer, because at some price point it stops being a decision and starts being a reflex. The same thing may be happening with intelligence right now. Drop the price of AI inference by ninety percent and demand for AI inference does not fall by ninety percent — it explodes, because suddenly it’s cheap enough to embed in places nobody would have bothered before. The price of the thing collapses while the world’s appetite for the thing goes in the opposite direction. Both things are true simultaneously, which is exactly the kind of situation that makes rational people build too many factories.

Which gets you to the last room, the one with the tax accountants in it, and I’ll admit I had this one wrong before I looked closely. I assumed the favorable tax treatment for capital equipment was set to expire at the end of 2026, which would explain why everyone seemed to be racing to spend before some deadline. It isn’t expiring. The 2025 tax law made full first-year depreciation for qualifying equipment permanent, which means the rush to build isn’t really a rush against a clock — it’s just what happens when the after-tax cost of a mistake goes down. Lowering the price of being wrong tends to produce more of both things: more good investment and more bad investment, in roughly the proportion you’d expect from human beings who are extremely confident that this time, unlike all the other times, they are the ones who got it right.

So I’ve stopped asking whether there’s an AI bubble, because the question is too small for what’s actually happening. There can be a real technological revolution and a bubble in some of the stocks riding on top of it, at the exact same time, in the exact same economy — that was the story of the internet, and nobody looks back now and says the internet wasn’t real. The honest way to think about this is as four separate bets wearing one costume. Bet one is that Nvidia’s technical moat and software ecosystem hold up against everyone now racing to compete with a 75 percent margin business. Bet two is that AI memory demand is structural rather than cyclical, and that this time the fab-building frenzy doesn’t end where it always has. Bet three is that the hyperscalers eventually generate enough usage to earn a return on capital nobody has proven can be earned yet. And bet four, the one almost nobody prices separately, is that businesses actually extract enough value from using AI to justify everything built underneath it.

Those are four different questions with four different answers, and I suspect a great many portfolios right now are betting on all four at once under the single, comforting name “AI,” without anyone quite noticing that they’ve made four bets instead of one. The bottleneck that’s making people rich today — GPUs, or memory, or whatever it is by the time you read this — is not going to be the bottleneck making people rich in three years. It never is. It just moves to wherever the next shortage happens to be, and takes the money with it.

I keep coming back to that sold-out warehouse. Somewhere out there is the shipment that finally isn’t sold out. Nobody rings a bell when it arrives.

Categories
AI Business

The Gravity of Compute

We are currently witnessing the single largest deployment of capital in human history. The “Hyperscalers”—the titans of our digital age—are pouring hundreds of billions of dollars into the ground, turning cash into concrete, copper, and silicon.

The prevailing narrative is one of unceasing, exponential growth: bigger models require bigger clusters, which require more power plants, which require more land. It relies on the assumption that the demand for centralized intelligence is insatiable and that the current architecture is the only way to feed it.

But history suggests that technology rarely moves in a straight line; it swings like a pendulum. Two forces are emerging from the periphery that could impact the ROI of this massive infrastructure build-out. One is hiding in your pocket, and the other is waiting in the sky.

A recent conversation with Gavin Baker outlines a potential “bear case” for datacenter compute demand: the rise of Edge AI.

We often assume we need the “God models”—the omniscient, trillion-parameter giants hosted in the cloud—for every interaction. But do we?

Baker suggests that within three years, our phones will possess the DRAM and battery density to run pruned versions of advanced models (like a Gemini 5 or Grok 4) locally. He paints a picture of a device capable of delivering 30 to 60 tokens per second at an “IQ of 115.”

“If that happens, if like 30 to 60 tokens at… a 115 IQ is good enough. I think that’s a bear case.” — Gavin Baker

Consider the implications of that specific number. An IQ of 115 isn’t omniscient, but it is competent. It is capable, nuanced, and helpful.

If Apple’s strategy succeeds—making the phone the primary distributor of privacy-safe, free, local intelligence—the vast majority of our daily queries will never leave the device. We will only reach for the cloud’s “God models” when we are truly stumped, much like we might consult a specialist only after our general practitioner has reached their limit. If 80% of inference happens on the edge for free, the economic model of the trillion-dollar data center begins to look fragile.

Then there is the second threat, one that attacks the terrestrial constraints of the data center itself: the Orbital Data Center. Elon Musk and SpaceX – along with Google’s Project Suncatcher – envision a future where the heavy lifting isn’t done on land, but in orbit. Space offers two things that are scarce and expensive on Earth: unlimited solar energy and an infinite heat sink for radiative cooling. If Starship can reliably loft “server racks” into orbit, the terrestrial moat of land and power grid access—currently the Hyperscalers’ greatest defensive asset—evaporates.

We are left with a fascinating juxtaposition. On one hand, we have the “Edge,” pulling intelligence down from the clouds and putting it into our hands, making it personal, private, and free. On the other, we have “Orbit,” threatening to lift the remaining heavy compute off the planet entirely to bypass the energy bottleneck.

There are hundreds of billions of dollars betting on a future of heavy, centralized gravity. But if the edge gets smart enough, and the orbit gets cheap enough, the gravity may have shifted.