Categories
AI

The Tempo of the Brake Pedal

Martin Casado said something this week that snapped a few loose thoughts into place. The frontier labs missed Jev, he argued, because they are building beings that speak. Software needed a model that chooses.

That reads like a product distinction. It is really a habit of mind.

For a few years we have taken a machine born in chat — text in, text out — and tried to cram it into ordinary programs. Parse the JSON. Retry when it rambles. Plead with it to stay inside the schema. Hope it does not invent a field. Casado’s word for this was right: janky.

Jev takes the other fork. Don’t generate the paragraph. Give the model a state and a menu. Let it pick, score, or answer yes or no, with a probability attached, in one pass, cheap enough and fast enough to live inside a loop. Train it for that job instead of training it to sound like someone worth talking to. The name winks at Jevons: make the unit of intelligence cheap enough and software will call it constantly. That is not a chatbot. That is a function.

I keep finding the same fork elsewhere.

Full self-driving is not AGI on wheels. AGI is the being. Driving is the chooser. The car does not need a meditation on the yellow light. It needs brake, hold, go, at a tempo no essay can match. Tesla and Waymo are not succeeding or failing as philosophers. They are succeeding or failing as systems that pick from a small set of actions in a messy world, millions of times a day.

Humanoid robots make the same confusion look even more like the story we want it to be. Two legs, two hands, tidy the kitchen. The demo is speech and generality. The work underneath is neither: a high-frequency controller choosing stance and contact, and a slower chooser deciding which mug, which drawer, abort or continue. Language is a compiler of intent. It is a poor motor cortex.

We blur the layers because the being is the story we want to tell. The chooser is the product that has to work.

There is a flattering counterargument: wait for the general model, and it will absorb the narrow jobs. Sometimes it will. But a lot of software’s value does not live there. It lives in the call that has to be right, cheap, and on time — route the ticket, score the lead, decide the grasp, stay in the lane. A being is optimized to continue. A chooser is optimized to decide. Different losses. Different machines.

None of this makes the labs foolish. If you are trying to build God, God speaks in natural language. That is a coherent ambition. It is just not the same ambition as making ordinary software reliable. We spent a few years pretending those were the same project. Jev is a reminder that they never were.

Labs build speakers. Software, cars, and robots need choosers. Speech is a wonderful interface and a miserable control loop.

I have been as taken as anyone by the being — the voice, the fluency, the sense that something is in there. Curiosity pulls that way. So does the stage. But the quieter question, the one Casado is actually pointing at, is whether most of the work was ever conversation at all. Maybe it was always a menu, a state, and a choice — and we were paying a novelist to raise his hand.

Somewhere right now a self-driving car is holding at a yellow light, saying nothing at all, deciding everything.

Categories
AI Infrastructure

The Weight of What’s Inside

Gigatexas. I watched the footage early yesterday, still in bed, before I was fully awake enough to know why I couldn’t stop. Not any single building — the simultaneity of it. Steel skeleton rising on the North Campus, where a dedicated line will eventually try to build ten million humanoid robots a year. An advanced chip fabrication building going up close enough to share a fence line with it, because the AI hardware and the AI bodies have apparently become the same argument. And underneath all of it, still running, still shipping, the original Model Y line that paid for everything else. Three or four enormous bets, at three or four different stages of doubt, on the same 2,100 acres, none of them waiting for the others to finish.

And then the second thought arrived, quieter than the first: this is the outside. A drone at four hundred feet can show you steel and concrete and rows of finished cars. It cannot show you the tooling, the calibration, the thousand small decisions about how a robot learns to close its hand around an object it has never held before. We were watching a shell form around something we couldn’t see into, and it would be easy to mistake the shell for the thing.

I started the Sarah Guo interview about an hour later, same morning, footage still fresh, and the two things turned out to be the same essay, just told in different registers — hers in argument, Gigatexas’s in steel.

Categories
AI

The Arithmetic of the Sold-Out Warehouse

In the spring of 2026, Nvidia reported a quarter in which it sold $81.6 billion worth of chips, wrote it up at a gross margin of nearly 75 percent, and casually mentioned that cloud GPUs were sold out. Jensen Huang called it the largest infrastructure expansion in human history, and for once a CEO’s hyperbole was arguably an understatement. Revenue was up 85 percent from a year earlier. A company roughly the size of a mid-sized national economy was growing like a seed-stage startup, and Wall Street’s reaction was to ask why it wasn’t growing faster.

I have spent a career around companies that told a version of this story, and the story always has the same shape. Something becomes scarce. Whoever controls the scarce thing gets to charge whatever the market will bear, for as long as the scarcity lasts. The interesting question was never whether Nvidia’s chips were good. Everyone agreed they were good. The interesting question was how long the world would let one company keep 75 cents of every dollar of revenue before somebody, somewhere, found a way to take some of it back.

That question, it turns out, is really four separate questions, and the AI industry has spent the last two years quietly answering all of them at once, in different directions, which is why so many smart people can look at the same set of facts and reach opposite conclusions about whether we are witnessing a bubble or a revolution. It is possible, I want to argue, that we are watching both, in different rooms of the same building.

Start with the money. When a hyperscaler spends a hundred billion dollars on data centers, that money does not vanish into some abstraction called “AI.” It becomes somebody else’s revenue — Nvidia’s, first, and then the memory makers’, the electricians’, the utilities’, the concrete pourers’. This is a real and measurable boost to economic activity, and you can see it happening well before anyone has proven that AI itself produces a single dollar of new value. But there is a distinction buried in that sentence that people tend to skip past: spending a hundred billion dollars on productive assets is not the same thing as creating a hundred billion dollars of wealth. The assets still have to earn their keep. Somebody has to use them for something worth more than they cost.

Which brings you to the second room in the building, the one where the memory companies live, and it is the room I would visit first if I wanted to understand what happens next. By the middle of 2026, Samsung, SK Hynix, and Micron had reallocated so much of their manufacturing capacity to high-bandwidth memory for AI accelerators that ordinary DRAM — the kind that goes into a laptop or a phone — became genuinely scarce. Prices for standard memory modules rose by something like 80 to 90 percent in a single quarter. SK Hynix posted an operating margin north of 70 percent. Micron’s profit rose more than sevenfold year over year. Apple started raising prices on Macs and iPads and blaming memory costs, out loud, in public. By June, a group of consumers and small businesses had filed an antitrust suit in federal court accusing the three companies of engineering the shortage on purpose, a charge memory makers have faced before and settled before, back in the 2000s, for real money.

I don’t know whether that lawsuit has merit. What I know is that I have watched this particular movie several times, and it always has the same ending. Scarcity produces extraordinary margins. Extraordinary margins summon capital. Capital builds capacity. Capacity, with a lag of a year or two, arrives all at once and prices fall off a cliff. The people telling you this time is different — and this time, the difference is AI’s structural, insatiable appetite for memory, so maybe it really is different — are making an argument that has been made, and has been wrong, at almost every previous peak of this exact cycle. Building a new fab takes eighteen to twenty-four months. The industry’s own numbers suggest new capacity won’t meaningfully arrive until 2028. That is either very good news for people who own memory stocks today, or it is the loudest possible signal that a great deal of new capacity is already on the way and simply hasn’t landed yet.

Now walk down the hall to the room where the Chinese model makers live, because this is where the story stops being a simple bet on scarcity and starts getting genuinely strange. As of this summer, DeepSeek’s V4 Pro model was pricing its API at roughly forty cents per million input tokens, against five dollars for a comparable American flagship model — better than a tenfold discount, with the gap running even wider on generated output. Alibaba’s Qwen and Moonshot’s Kimi were sitting in a similar band. Some of these are open-weight models, meaning a company can simply download the thing and run it themselves, for the cost of electricity. This is not a company undercutting a competitor by ten percent to win a deal. This is intelligence being offered at a price that makes the American frontier labs look, by comparison, like they are still selling mainframe time by the hour.

If you take that seriously, it forces an uncomfortable question. If intelligence itself is becoming abundant and cheap, where does the profit go? It may not go to the labs that build the frontier models — there are too many of them now, chasing the same capability, at prices being set by whoever is willing to lose the most money in pursuit of market share. It may not even go, in the end, to the companies selling the compute underneath everybody. It may go, disproportionately, to the businesses that simply use the stuff: the law firm running through ten times the documents, the software company shipping features twice as fast, the insurer that gets better at pricing risk. Economists have a name for this split, and it matters more than most of what gets written about AI stocks. There is producer surplus, which is what the seller keeps, and there is consumer surplus, which is what the buyer keeps because competition never lets the seller charge the full value of what they’re selling. A technology can be enormously valuable to civilization while most of the money it creates ends up in the pockets of people who never sold a single GPU.

Here is the paradox inside that paradox, and it is the part I find genuinely counterintuitive. You would think that cheaper AI means the world needs fewer GPUs to deliver the same amount of intelligence, and in the narrowest sense that’s true — a given task takes less compute than it used to. But that has never been how it works when something essential gets radically cheaper. Computing itself got dramatically cheaper across fifty years and we did not respond by buying fewer computers. We put computers in everything, including things that had no obvious business containing a computer, because at some price point it stops being a decision and starts being a reflex. The same thing may be happening with intelligence right now. Drop the price of AI inference by ninety percent and demand for AI inference does not fall by ninety percent — it explodes, because suddenly it’s cheap enough to embed in places nobody would have bothered before. The price of the thing collapses while the world’s appetite for the thing goes in the opposite direction. Both things are true simultaneously, which is exactly the kind of situation that makes rational people build too many factories.

Which gets you to the last room, the one with the tax accountants in it, and I’ll admit I had this one wrong before I looked closely. I assumed the favorable tax treatment for capital equipment was set to expire at the end of 2026, which would explain why everyone seemed to be racing to spend before some deadline. It isn’t expiring. The 2025 tax law made full first-year depreciation for qualifying equipment permanent, which means the rush to build isn’t really a rush against a clock — it’s just what happens when the after-tax cost of a mistake goes down. Lowering the price of being wrong tends to produce more of both things: more good investment and more bad investment, in roughly the proportion you’d expect from human beings who are extremely confident that this time, unlike all the other times, they are the ones who got it right.

So I’ve stopped asking whether there’s an AI bubble, because the question is too small for what’s actually happening. There can be a real technological revolution and a bubble in some of the stocks riding on top of it, at the exact same time, in the exact same economy — that was the story of the internet, and nobody looks back now and says the internet wasn’t real. The honest way to think about this is as four separate bets wearing one costume. Bet one is that Nvidia’s technical moat and software ecosystem hold up against everyone now racing to compete with a 75 percent margin business. Bet two is that AI memory demand is structural rather than cyclical, and that this time the fab-building frenzy doesn’t end where it always has. Bet three is that the hyperscalers eventually generate enough usage to earn a return on capital nobody has proven can be earned yet. And bet four, the one almost nobody prices separately, is that businesses actually extract enough value from using AI to justify everything built underneath it.

Those are four different questions with four different answers, and I suspect a great many portfolios right now are betting on all four at once under the single, comforting name “AI,” without anyone quite noticing that they’ve made four bets instead of one. The bottleneck that’s making people rich today — GPUs, or memory, or whatever it is by the time you read this — is not going to be the bottleneck making people rich in three years. It never is. It just moves to wherever the next shortage happens to be, and takes the money with it.

I keep coming back to that sold-out warehouse. Somewhere out there is the shipment that finally isn’t sold out. Nobody rings a bell when it arrives.