Categories
AI

The Kitchen, Not the Farm

There is a sentence buried in Thinking Machines Lab’s release notes for Inkling, its first proprietary model, that most companies would never let out the door. Describing their own creation, the company states plainly that Inkling is “not the strongest overall model available today, open or closed.”

Read that again. A startup that raised two billion dollars in seed funding at a twelve-billion-dollar valuation, founded by OpenAI’s former CTO and staffed with veterans of the labs currently locked in the most capital-intensive arms race in corporate history, shipped its debut model with an admission of inferiority attached to the label. Not buried in a footnote. Stated in the announcement.

It seems like this was the clearest signal yet that the frontier-capability race may be the wrong game, and that durable value in enterprise AI accrues not to whoever has the smartest model, but to whoever owns the layer where that model gets adapted to a particular customer’s purpose.

As I’ve thought about it, the AI industry seems to be stratifying into three distinct businesses, each with different economics, occupied by a different cast of companies.

It begins with the farm, where the raw ingredients get grown. Then there’s the kitchen, the capital equipment that makes skilled cooking possible at scale. Lastly there’s the restaurant, where somebody who understands a specific customer takes the ingredients, uses the kitchen, and puts a particular dish in front of a particular diner who is paying for a complete meal, not just the flour or the vegetables. Thinking Machines seems to me like a clear example of a company trying to explain which of those businesses it’s actually in. It is not the only one.

The sequence, read backward

Founded in February 2025. Silent for over a year. Then, last October, the company’s first product emerged โ€” and it wasn’t a chatbot, wasn’t an assistant, wasn’t anything a consumer like me would recognize or understand. It was Tinker, a fine-tuning API. Infrastructure for customizing other people’s models, shipped before the company had released a model of its own.

That sequencing is the tell. A company chasing frontier supremacy builds the model first and the tooling around it later, the way the frontier AI labs have all done. Thinking Machines inverted the order. It built the workshop before it built anything to put in the workshop window, which only makes sense if the workshop was always the product.

Inkling, released this month, doesn’t reverse that logic. It completes it. The model is described in the company’s own materials as “an extremely knowledgeable, generalist base that can be extended via fine-tuning” โ€” language that positions the model itself as raw material, not a finished good. It ships with full open weights, day-zero availability on Tinker, and a name chosen, according to the company, to evoke “an idea in its earliest stage, with the potential to grow into something greater.” Even the naming is a thesis statement. Inkling is not meant to be the only thing you use. It’s meant to be the thing you start from.

In the farm-kitchen-restaurant frame, it seems like Thinking Machines is trying to own two levels of the stack at once. Inkling is the farm โ€” grown at real expense, forty-five trillion tokens of training data, frontier-scale compute. Tinker is the kitchen โ€” the induction range and the walk-in fridge, sold as a service to whoever wants to cook. What Thinking Machines has explicitly declined to be, by its own admission, is the restaurant. They are not trying to serve you the best possible dish. They are trying to make sure that whoever does serve you that dish is buying their ingredients and standing at their stove and cooking in their kitchen.

The manifesto that preceded the model

A company doesn’t back into a strategy this coherent by accident. Earlier this month โ€” before Inkling shipped โ€” the lab published a position paper arguing that most AI today is trained in a handful of places and then frozen, a design that by its nature excludes the people the model is meant to serve. Their proposed alternative: AI that is distributed, customizable, and shaped by the people using it, not the lab that built it.

Mira Murati has said the same thing more plainly, and said it a year before Inkling existed, back when Tinker launched. Her framing wasn’t about building the smartest model. It was about making “frontier capabilities much more accessible to all people” โ€” democratization as the mission, not capability supremacy. That is a genuinely different objective function than the one driving her former employer, and it was declared outright, not discovered after the fact to explain a disappointing benchmark result.

Inkling is a 975-billion-parameter mixture-of-experts model trained on forty-five trillion tokens across text, image, audio, and video, with a context window stretching to a million tokens. That is frontier-scale compute expenditure. This isn’t a company that ran out of runway and settled for a smaller ambition. It’s a company that spent frontier-level resources and then declined to spend the final increment chasing benchmark supremacy, presumably because the return on that increment doesn’t show up in the business they’re building.

A second detail: Inkling reportedly uses one-third the tokens of Nemotron 3 Ultra to hit equivalent performance on agentic coding benchmarks. That’s not a capability retreat โ€” that’s a capability choice, optimizing for efficiency and cost-per-task rather than raw benchmark position. And the company is previewing a smaller sibling model alongside Inkling, suggesting a family strategy across sizes rather than a single mid-tier release.

The restaurant next door: Palantir

Thinking Machines isn’t the only company making this bet โ€” and looking at who else is making it shows not everyone is occupying the same layer.

Earlier this month Palantir and Nvidia announced a “Sovereign AI Operating System” โ€” Nvidia’s open Nemotron models, fine-tuned on a customer’s own data, running on Nvidia hardware inside that customer’s own air-gapped network, with Palantir’s Ontology and Foundry software layered on top. CEO Alex Karp pointed out that his enterprise customers don’t want to risk sharing their IP with frontier model providers and asked simply why wouldn’t they control the weights?

It’s tempting to read this as the same argument Thinking Machines is making. It isn’t, quite. What Palantir is selling is the restaurant: the finished, seasoned, plated product โ€” an air-gapped AI system wired into a specific government agency’s or enterprise customer’s actual workflows, with “you control the weights” as the pitch that closes the deal. Palantir isn’t growing wheat. It’s the chef, working with ingredients somebody else grew. Somebody who could be trusted.

Another restaurant: Sierra

Sierra, Bret Taylor and Clay Bavor’s customer-support agent company, makes the same choice even more starkly. Sierra’s own technical writing describes a “constellation of models” architecture: rather than betting on a single LLM, Sierra routes each task inside a customer-service agent to whichever model โ€” from OpenAI, Anthropic, Meta, or elsewhere โ€” handles it best, and explicitly says it invests “in fine-tuned models where off-the-shelf models fail to meet our constraints.” Fine-tuning shows up in Sierra’s stack as one tool among several, alongside retrieval and layered “supervisor” models that catch mistakes before a customer sees them. Sierra has no interest in being a model company or an infrastructure company. It wants to be the restaurant that happens to keep a few specialty ingredients in the walk-in that nobody else stocks, because the dish needs them and they know just how to include them.

Mapping the rest of the stack

The farm-kitchen-restaurant split shows up everywhere the fine-tuning economy has organized itself.

The kitchen-builders โ€” companies selling fine-tuning infrastructure to whoever wants to cook with it, indifferent to what gets made โ€” now form a crowded field: Thinking Machines’ Tinker, Together AI, Fireworks AI, Predibase, OpenPipe, Baseten, Modal, Databricks’ Mosaic stack, and newer entrants like Nebius’s Token Factory and Prime Intellect. None of them care whether you’re building a coding agent, a legal research tool, or a customer-service bot.

The restaurants โ€” companies where fine-tuning is invisible plumbing inside a finished, vertical product โ€” include Palantir and Sierra, and many others. The addressable market for fine-tuning seems to include almost every possible enterprise adopting AI.

What’s seems unusual about Thinking Machines is that it’s trying to be the farm and the kitchen simultaneously while declining, by its own public admission, to be the restaurant. Most companies pick one layer and defend it. Thinking Machines is betting that owning two of the three is the more durable position โ€” grow the flour, own the stove, and let Palantir, Sierra, and a thousand enterprise engineering teams fight over who plates the dish.

The same stack, built by design

As I was thinking about this, I wondered how this relates to the AI activities underway in China. It seems that China’s AI industry maps onto this same three-layer structure with unusual clarity โ€” and one genuine wrinkle the American version doesn’t have.

The farm is crowded and innovating on a different axis than size: DeepSeek, Alibaba’s Qwen, Zhipu AI, Moonshot AI, MiniMax, ByteDance’s Doubao and Seedance. The standout isn’t scale, it’s efficiency โ€” DeepSeek’s V3.2 reportedly uses a novel sparse attention mechanism to nearly match GPT-5 and Gemini 3 on complex reasoning despite far less compute, a different kind of farming: not more wheat, but wheat bred to need less water. Qwen has become the default soil for the rest of the world’s kitchens, generating over 100,000 derivative fine-tunes on Hugging Face. VC’s in Silicon Valley note how frequently their startup companies are building on Qwen.

The kitchen layer has its own SiliconFlow โ€” a Beijing infrastructure startup, backed by Alibaba Cloud, that bills itself as the neutral layer between AI applications and hardware. It solves a problem others never had to: China’s compute runs across fragmented domestic chips, Huawei’s Ascend line chief among them, that don’t share Nvidia’s CUDA ecosystem. SiliconFlow abstracts that fragmentation away โ€” it became the fastest platform serving DeepSeek traffic, and the only large provider running DeepSeek on Ascend chips instead of Nvidia’s. That’s a stove engineered to burn whatever fuel is in the tank that week, a direct product of the U.S. chip export controls rather than any inherent technical edge. Volcano Engine, Alibaba Cloud’s PAI, and Baidu’s Qianfan are versions of the same layer.

The restaurant layer is where China’s picture diverges most from Palantir and Sierra’s venture-funded improvisation: it’s named industrial policy.

Beijing’s “AI+” initiative targets seventy percent sectoral AI penetration by 2027, ninety by 2030 โ€” fine-tuned vertical deployment treated the way past five-year plans treated high-speed rail. The players read like a sector directory: SenseTime for vision and embodied AI, iFlytek for speech in education and government, Baichuan Intelligence for healthcare, 4Paradigm for finance and industry, each fine-tuning a general base into something that only makes sense inside one workflow โ€” a hospital’s diagnostic support tool, a bank’s risk model, an industrial inspection line.

The bet

Every frontier lab is implicitly betting that intelligence is the scarce resource, and that whoever has the most of it wins the enterprise market by default. Thinking Machines, Palantir, Sierra, and many others are all, in their different ways, betting against that premise โ€” that raw intelligence is commoditizing faster than the frontier labs’ spending would suggest, and that the scarce resource has already migrated to whichever layer turns a generalist model into a specific customer’s model.

Thinking Machines is betting the moat moved to the farm-and-kitchen layer. Palantir, Sierra and others are betting it moved further still, to the restaurant, where nobody cares whose flour was used as long as the dish is right. China is betting on all three layers at once, with the state underwriting the bet directly.

It is a curious thing for me to watch companies with this much money and this much talent choose not to fight for the title of smartest model in the room. It is also a curious thing to watch them explain why, in public, in the first paragraph of an announcement.

But I think I’m beginning to understand.

Categories
AI

Memory Optimization

After thinking about the very high margins that Micron recently reported (greater than Nvidiaโ€™s) I asked Grok 4.5:

What startups are pursuing technology that reduces memory requirements for ai models and inference?

Several startups are actively developing technologies to slash memory requirements for AI models and inference. These efforts target the core bottlenecks in large language models (LLMs) and multimodal systemsโ€”primarily through extreme quantization (reducing bit precision of weights), alternative architectures that avoid quadratic memory scaling (e.g., state-space models or hybrids), sparsity/partial activation (like Mixture-of-Experts or neuron-level sparsity), efficient runtimes/kernels, and hardware innovations like in-memory computing that minimize data movement.

This space is especially active for on-device/edge AI (phones, laptops, embedded systems) and cost/energy-efficient cloud inference, where fitting larger or more capable models into limited RAM/VRAM, reducing KV cache bloat in Transformers, or lowering bandwidth demands delivers major gains.

Here are some of the most prominent startups and their approaches (as of mid-2026):

Extreme Quantization & Model Compression

  • PrismML (Caltech spinout, backed by Khosla Ventures): Focuses on extreme low-bit architecturesโ€”1-bit binary (โ€œBonsaiโ€) and ternary (1.58-bit, weights in {-1, 0, +1}). They compressed Alibabaโ€™s Qwen3.6 27B model from ~54 GB (FP16) to under 4 GB and ran the full dense 27B model on an iPhone 17 Pro. Claims include up to 14ร— smaller memory footprint, 8ร— faster inference, and significantly lower energy use, with competitive or better benchmark performance. They have open-sourced Bonsai models (including smaller 8B/4B/1.7B variants) under Apache 2.0 and are in discussions with Apple. This represents one of the most aggressive commercial pushes into 1-bit/ternary models for on-device deployment.
  • Mobius Labs (Berlin): Developed Half-Quadratic Quantization (HQQ), a fast, calibration-light post-training quantization method that enables high-accuracy low-bit models (including aggressive 2-4 bit). They demonstrated quantizing Llama 70B to run on a single GPU instead of four without major accuracy loss, directly cutting memory and compute needs. Their work extends to FP4 optimizations and integrates with frameworks like vLLM.
Alternative Architectures for Inherent Memory Efficiency
  • Liquid AI (MIT spinoff): Builds Liquid Foundation Models (LFM / LFM2 series)โ€”hybrid architectures combining gated short convolutions with grouped-query attention (GQA) blocks, plus MoE variants. These deliver substantially lower memory footprints than Transformers (especially for long contexts, avoiding massive KV cache growth), faster prefill/decode (up to 2ร— on CPU in some cases), and strong on-device performance. Examples include tiny models (230Mโ€“350M params, often
  • Cartesia: Specializes in state-space models (SSMs) and Mamba-derived/hybrid architectures. These offer near-linear (or better) scaling with sequence length and near-constant memory complexity during inference (fixed or minimal growth vs. quadratic KV cache in Transformers). This enables efficient long-context processing on smaller devices with lower energy use. They provide on-device libraries (e.g., โ€œEdgeโ€) and models like hybrid Mamba-2 variants optimized for throughput and fixed memory footprints.
On-Device/Edge Runtimes & Optimizations
  • Cactus (Y Combinator-backed): Builds cross-platform, energy-efficient inference runtimes and kernels for mobile, wearables, and low-power devices. Supports aggressive quantization (down to 2-bit), delivers very low latency (sub-50 ms time-to-first-token), and keeps models small (e.g., hundreds of MB for small models). Emphasizes privacy (fully local) and broad model support (Llama, Qwen, Gemma, etc.).
  • Embedl: Focuses on edge-optimized multimodal models and mixed-precision quantization (e.g., W4A16-Edge2 with selective FP16 layers + โ€œFlashHeadโ€ optimizations). Enables fast multimodal reasoning on devices with under 8 GB RAM (e.g., Jetson Orin Nano/AGX), with significant speedups over baselines while recovering most accuracy.
Hardware Innovations Reducing Memory Pressure
  • d-Matrix: Develops digital in-memory computing (DIMC / 3DIMC) chips and platforms (e.g., Corsair) that integrate compute directly with/near memory (initially SRAM, now advancing to 3D DRAM). This drastically cuts data movementโ€”the dominant memory bandwidth and energy bottleneck in inference. Claims include major speedups (up to 10ร— in some workloads) and better efficiency vs. traditional GPU/HBM setups, enabling higher performance with more manageable overall memory system requirements.
  • Hailo: Designs efficient edge AI inference chips (e.g., Hailo-10H series) optimized for low-power LLM and multimodal workloads. Their chips target 1โ€“3B parameter models at very low power envelopes (~2โ€“2.5 W), making deployment feasible on memory- and power-constrained edge devices where traditional hardware would struggle.
Other Notable Mentions
  • Tiiny AI: Uses neuron-level sparse activation (โ€œTurboSparseโ€) + heterogeneous engines for running very large models (up to 120B) fully on-device/offline in compact hardware.
  • Broader ecosystem players (inference platforms like Together AI or Fireworks) heavily optimize serving (quantization, caching, KV cache management) but are more infrastructure-focused than core technology innovators in compression/architecture.
Key Trends and Context

Quantization (especially post-training or quantization-aware) remains the most immediate lever and is widely adopted. Architectural shifts (SSMs, hybrids, MoE with low active parameters, convolutions) address fundamental scaling issues like KV cache memory growth. Hardware approaches (in-memory compute, high on-chip SRAM, efficient NPUs) tackle bandwidth and data movement. Many solutions combine these (e.g., architecture + quantization + sparsity).

Progress is rapid, driven by on-device AI demands (Appleโ€™s interest in PrismML is notable) and the economics of scaling inference. Many models are open-sourced or have open weights, accelerating adoption. Results vary by workloadโ€”extreme 1-bit models excel in certain scenarios but may need custom kernels; SSM/hybrid models shine for long context or linear scaling needs.

For the absolute latest benchmarks or specific model deployments, check company sites (e.g., PrismMLโ€™s Bonsai releases, Liquid AIโ€™s LFM models on Hugging Face) or recent evaluations, as the field evolves quickly.

Interesting!

Categories
AI Semiconductors

The Margin of the Weather

A company that has sold memory chips for forty years โ€” memory, one of the most humiliatingly commoditized products in capitalism, a business that has bankrupted entire Korean and Japanese conglomerates teaching each other lessons about discipline โ€” is about to make more money in twelve months than in the previous four decades combined.

Samsung’s chip chief told a room of his own employees: this year’s profit will exceed everything the division has earned since the 1970s. Forty years of grinding, erased by one fiscal year. You’d think they’d invented something.

They hadn’t. Everyone building an AI data center needs memory. Nobody built enough factories. Samsung was one of three companies on earth able to supply the shortfall, and the price of a chip that costs what it always cost went up fifty percent. Samsung kept the difference. Not innovation. What happens to a farmer when the drought hits every field but his.

We don’t credit the lucky farmer with genius. We say: good year. And we don’t expect the good year to repeat. Rain comes back. The price falls. Scarcity is weather, not a personality trait.

There’s a real achievement in this story too, and it has nothing to do with the weather. A year ago Samsung failed to qualify its most advanced memory for Nvidia’s systems โ€” performance problems, a rival getting the business instead. The engineers went back and fixed it. That’s the actual skill in this company’s year: unglamorous, uncelebrated at the town hall, worth nothing next to the number that got the confetti. The competence arrived quietly, on a different chip, in a different meeting, and nobody’s putting that on a plaque.

The stock market didn’t put it on one either, but it seemed to know the difference. Best quarter in Samsung’s history โ€” profit nineteen times the year before โ€” and the shares fell seven percent. Not despite the earnings. The gain had already been priced in, the shares having run up a hundred and fifty percent on the expectation of exactly this number, so the number’s arrival became a ceiling instead of a floor. A market rewards discovery. It does not reward weather. Had investors believed Samsung built something durable โ€” the Nvidia qualification, the years of engineering behind it โ€” the stock would have ripped, the way See’s Candies or Apple gets rewarded quarter after quarter, because everyone agrees the thing generating the money isn’t going anywhere. Instead the market glanced at the record harvest and asked, politely, whether it would rain again next year.

Analysts insist the shortage holds through next year. Someone always insists that, right before it doesn’t. Fabs get built. Capacity catches the demand that summoned it, the way it always has, and the cycle ends the way memory cycles end โ€” too much supply chasing too little demand, margins reverting toward the number they were always going to revert toward. Nobody knows if this time is different. A company just posted the best year of its life, on a windfall it didn’t earn and a fix it did, and the market โ€” which has seen droughts end before โ€” hasn’t decided yet which one it’s watching.

Categories
AI Apple Google

The Floor

I compared the frontier to a three-star chef making grilled cheese in “Context Rot” โ€” the smartest models on earth spending most of their time on work beneath them, the way a chef trained at Le Bernardin might still melt cheese between two slices of bread on a Tuesday night and call it dinner. The comfort was the point: if the sharpest tool is saved for hard problems and something merely-very-good handles the rest, nobody’s losing anything. The floor was never the interesting part.

I’ve kept turning the joke over, and I think I had the wrong worry.

Watch what companies do with their AI spend, not what they say. Coinbase moved engineers off frontier models onto open weights and cut its AI spend nearly in half while usage kept climbing. Nvidia runs a closed model as orchestrator and routes the actual volume โ€” the daily uncelebrated bulk of it โ€” to open weights it controls. The frontier is becoming a dispatcher, deciding where the request goes and rarely doing the work itself. The instinct is to worry about whose open weights end up running that volume, and right now the most capable ones at scale are Chinese โ€” GLM, Kimi โ€” which makes it tempting to read this as a contest America is quietly losing: the floor of the AI economy built somewhere else, at a price export controls can’t touch. You cannot embargo a file already downloaded. You cannot price-match free.

But that framing has a hole. Google’s own Gemma family is open-weight and good enough to handle that daily volume without anyone reaching for GLM or Kimi. “Open weights are a Chinese story” only holds if you don’t count the open models the company running Android and half the internet’s search traffic has already shipped.

And once I saw that hole, a bigger one opened behind it. I’ve been trying Apple’s new Siri โ€” arriving with iOS 27 this fall, genuinely surprisingly good in beta โ€” and it made me realize open weights, of any nationality, were never going to cook most of the world’s dinners. Apple and Google are.

Consider what actually determines where the world’s routine inference runs. Not which model benchmarks best, not which weights are downloadable โ€” what’s already installed. Apple ships to well over a billion active devices before routing a single query through Siri’s new architecture. Nobody has to be persuaded to try it, or hear about it on a podcast; it’s the thing that answers when you press the button you’ve pressed for a decade. Google owns the search bar and the Android default the same way. Between them, that’s most of the world’s phones โ€” and phones are where most of the world’s questions get asked.

The open-weight framing assumes the floor is up for grabs, that whoever ships the best free model wins the daily grind by merit. But the floor was never a bazaar. It’s a set of defaults, owned by whoever already has the device in your hand, not whoever holds the most generous license. Apple didn’t need to win the model war to win this. Its heaviest reasoning tier is built with Google, running on Nvidia chips in Google’s cloud, under a deal reported at roughly a billion dollars a year โ€” Apple doesn’t fully own the engine doing the thinking. It doesn’t need to. It owns the button.

That’s a quieter concentration than an export-controls fight, and a harder one to dislodge. An open model can be forked, distilled, undercut, or out-competed by the next release. A billion phones with an assistant built into the lock screen cannot be routed around. Whoever’s weights hum underneath barely matters, the way it barely matters to a diner which supplier delivered the flour. What matters is whose kitchen the meal came from, and whose name is on the door.

The grilled-cheese chef was never the risk. Two chefs are about to own nearly every kitchen on earth, and most of us will never notice โ€” because a kitchen you’ve been eating out of for a decade doesn’t feel like something that was won. It just feels like home.

Owning the kitchen and getting paid for what’s cooked in it, though, turn out to be two different questions. That one’s for another post.

Categories
Aging AI Business Living

The Being Phase

There is a metric making the rounds in technology investing circles that is, on its face, about market share and revenue concentration. Alex Sacerdote of Whale Rock Capital calls it the New Rule of 40 for AI. The formula is simple: take the percentage of a companyโ€™s sales derived from AI, add its percentage market share in that AI category, and if the sum reaches 40, you have a winner. Celestica, a company most people have never heard of, scores extraordinarily well. It owns somewhere between half and sixty percent of the cloud Ethernet white-box switch market. NVIDIA doesnโ€™t need a formula. It simply is what it is.

Sacerdote designed the metric to cut through a specific kind of noise โ€” the companies claiming AI exposure they donโ€™t actually have, the giants whose AI revenue hovers at one or two percent of their base while their press releases suggest otherwise. The framework is a detector. It finds the companies that have stopped becoming AI infrastructure and started simply being it.

I found myself less interested in the companies than in that distinction.


I spent years at Visa watching a network that had long since crossed that threshold. By the time I arrived, Visa wasnโ€™t becoming the global payments infrastructure. It was the global payments infrastructure. The work was real โ€” fraud detection, modeling, the daily labor of keeping something enormous running โ€” but the existential question had been settled before I got there. The network existed. Merchants accepted it because cardholders carried it. Cardholders carried it because merchants accepted it. That loop had been closing for decades. We were custodians of a fait accompli.

Thereโ€™s a particular feeling to working inside something that has already won. Itโ€™s not complacency exactly. The problems are genuine and the stakes are high. But the uncertainty has a different quality โ€” itโ€™s operational uncertainty, not existential uncertainty. Youโ€™re not asking whether the thing will survive. Youโ€™re asking how to run it well.

I didnโ€™t have language for that distinction then. Sacerdoteโ€™s metric gives me some. The companies that score highest on his New Rule of 40 have resolved their existential question. Theyโ€™re not fighting for position. Theyโ€™re administering a position already held.


The question that has followed me out of that career, and out of several decades of watching technology cycles turn, is simpler and more personal than any investment framework.

When did I cross that line myself?


I have been writing at sjl.us since 2001. Thatโ€™s not a boast โ€” itโ€™s a data point. Twenty-five years of thinking out loud, of ideas arriving rather than being argued, of the specific memory as structural anchor. The blog is not becoming anything. It is what it is: a record of a mind moving through time, accumulated into something that has its own weight and shape.

The book on payments systems exists. The career at Visa exists. The photographs exist. The train journeys exist. The years in Dayton exist, and the years on the Peninsula, and the particular way the light falls on the California coast at Pescadero in the late afternoon โ€” when the fog is still offshore and the hills are improbably green and everything goes briefly, completely quiet, as if the world is deciding whether to continue.

These are not things I am building toward. They are things I am.

Sacerdote would say I have high market share in a specific category. The category is small โ€” one person, one particular configuration of experience and attention and accumulated knowing โ€” but the share is essentially total. There is no competitor for the position of having lived this particular life. The moat is absolute. The switching costs are infinite.

I used to find that thought melancholy. The narrowing as loss. The aperture closing on what remains.

Iโ€™m not sure I find it melancholy anymore.


The L-Curve, Sacerdote says, is a long flatline followed by a vertical explosion. The tinkering phase, then the moment of lift. He means it as a description of demand curves for technology infrastructure. But I recognize the shape from somewhere closer. The long middle of a life, building and becoming, and then the morning you wake up and realize the building is substantially done. What remains is the being.

Thatโ€™s not an ending. Itโ€™s a different kind of beginning.


Sacerdoteโ€™s metric will eventually stop working. All frameworks do. The AI infrastructure cycle will mature, the L-Curves will flatten, and some new measure will emerge to find the next thing that is just beginning to become what it will be. Thatโ€™s the nature of markets. The detector has to change as the signal changes.

But thereโ€™s a complication worth naming. Analysts at Citadel Securities published a note recently observing that even the most powerful technologies must pass through the prosaic discipline of cost curves, capacity constraints, and marginal returns. Token bills are arriving unexpectedly. Compute is scarce. The vision of AI as ubiquitous, frictionless, and immediate is colliding with physical reality. Their conclusion: asset prices will periodically be forced to reconcile ambition with physical constraint.

Thatโ€™s not a refutation of Sacerdote. Itโ€™s a reminder that feeling like youโ€™ve arrived and having actually arrived are different things. The being phase has to be load-tested. The position has to hold under pressure.

I think about the fiber optics Corning is laying into the massive data center clusters โ€” ultra-thin, bendable, carrying more light than anything that came before. The cable doesnโ€™t know itโ€™s infrastructure. It just carries what itโ€™s given, at the speed itโ€™s capable of, across whatever distance is required. It doesnโ€™t matter what the cable believes about itself. What matters is whether the light actually moves.

That seems right to me. You become what you are over a long time, largely without noticing. And then one day someone builds a metric that accidentally describes your life, and you recognize yourself in it, and you think: yes. Thatโ€™s the shape of it. High concentration. High share. A moat that deepened while you were looking elsewhere.

But the moat still has to hold.

The being phase, it turns out, is not the end of something. Itโ€™s the proof that something was built. And the daily question โ€” for companies, for infrastructure, for a person in his late seventies still writing, still paying attention โ€” is whether what was built is actually load-bearing.

You donโ€™t get to stop finding out.

Categories
Business History IBM Infrastructure Nvidia Programming Semiconductors

The Half-Life of Moats

Prompted by an article on X by @magicsilicon on the CUDA moat. Research and drafting assistance from my AI intern assistant Clark.

The NVIDIA H100 looks, in retrospect, like an inevitability. It wasnโ€™t.

What Jensen Huang built is more accurately understood as a sixteen-year accumulation of optionality โ€” a platform investment made in 2006 for a market that wouldnโ€™t fully materialize until 2022. NVIDIA intros the G80 architecture in November 2006, laying the groundwork for CUDAโ€™s release a few months later. The stated ambition was to let scientists write C++ that ran on GPU cores without needing to understand 3D graphics pipelines. The unstated bet was that parallel computation would eventually matter for something bigger than rendering shadows in video games.

For sixteen years, it mostly didnโ€™t. Not at scale. Not commercially. CUDA lived in research labs and HPC clusters. It attracted a small, devoted, and economically marginal user base โ€” the kind that papers cite but investors ignore. NVIDIA kept investing in it anyway: cuDNN for deep learning operations, cuBLAS for linear algebra, a layered ecosystem of libraries that made CUDA not just accessible but nearly irreplaceable for anyone doing serious numerical computation. When TensorFlow and PyTorch emerged as the standard frameworks for neural network research, they didnโ€™t adopt CUDA because it was the only option. They adopted it because CUDA was where the optimized kernels already lived.

AlexNet won the ImageNet competition in 2012 and did it on two NVIDIA GPUs. The deep learning community noticed immediately. The financial community largely did not.

Then ChatGPT launched in November 2022, and suddenly everyone needed H100s they couldnโ€™t get.


The parallel to Intel is instructive and also undersells how strange this kind of story looks while youโ€™re living through it. Intel was founded in 1968 as a memory company. DRAM. The founders โ€” Noyce, Moore, Grove โ€” were materials scientists and engineers who believed the future was in silicon memory chips. They were right, briefly: in the early 1970s Intel dominated the DRAM market. By 1984, that share had collapsed to 1.3%, ceded almost entirely to Japanese manufacturers who had commoditized the product.

What saved Intel wasnโ€™t a pivot so much as a realization that a stopgap had become a foundation. The 8086, conceived in 1976 as an internal hedge and launched in 1978 was never supposed to matter. It was a 16-bit processor designed to hold off Zilog while Intel finished its ambitious 32-bit iAPX 432 architecture. The 8086 was assigned to a single engineer. โ€œIf management had any inkling that this architecture would live on through many generations,โ€ its designer Stephen Morse later recalled, โ€œthey never would have trusted this task to a single person.โ€

IBM chose the 8088 โ€” a cost-reduced variant โ€” for the original IBM PC in 1981. That decision wasnโ€™t destiny, it was simply a procurement. And yet from that accident of selection, Intelโ€™s x86 line became the backbone of personal computing for four decades. The Pentium in 1993 was Intelโ€™s Wintel moment โ€” the flag bearer the @magicsilicon tweet gestures at โ€” but the flag had been quietly sewn since 1978.


What these histories share is not just a pattern of โ€œslow build, explosive payoff.โ€ The structural similarity is subtler: in both cases, the moat was a software abstraction layer built on top of hardware. Intelโ€™s real lock-in wasnโ€™t transistor count or clock speed. It was backward compatibility โ€” the commitment, formalized with the 80386 in 1985, that every future Intel chip would run software written for older ones. That promise created a flywheel that trapped developers and buyers in a virtuous (for Intel) dependency loop for decades.

CUDA is the same architecture at a different layer. The lock-in isnโ€™t the H100โ€™s 80 gigabytes of HBM3. Itโ€™s that switching to an AMD MI300X or Google TPU means potentially rewriting training pipelines that have been optimized against CUDA kernels for years. AMDโ€™s ROCm platform exists. It is, by most accounts, maturing. Engineers who have tried the migration report that it costs months and hundreds of thousands of dollars. The moat isnโ€™t a wall. Itโ€™s accumulated friction โ€” the switching cost of a decade of engineering decisions baked into codebases that no one wants to touch.


But to find the actual origin of this pattern, you have to go back further than Intel. To 1964, and to a decision IBM made that Fred Brooks โ€” its project manager โ€” called a bet-the-business move.

The IBM System/360 was announced on April 7, 1964, after five years of turbulent internal development. What it introduced wasnโ€™t just a new computer. It was a new concept: the separation of architecture from implementation. Before the 360, IBM ran five incompatible product lines simultaneously. A customer who outgrew their machine had to scrap all existing software and start over. The 360 replaced all five lines with a single unified architecture โ€” six models covering a fiftyfold performance range, all running the same operating system, all sharing the same instruction set. The name itself encoded the ambition: 360 degrees, all directions, all users.

Gene Amdahl, the 360โ€™s chief architect, had a precise formulation for what this meant: the architecture was โ€œan interface for which software is written, independent of any implementation.โ€ The Principles of Operation manual described what the machine did; separate Functional Characteristics documents described how each model did it. This distinction โ€” separating the contract from the execution โ€” was genuinely new. Itโ€™s the conceptual root of everything that came after.

The 360 generated over $100 billion in revenue for IBM and established the first platform business model in computing. Jim Collins would later rank it alongside the Model T and the Boeing 707 as one of the three greatest business achievements of the twentieth century. But its deepest legacy was architectural: the insight that if you make your abstraction layer the standard, the hardware underneath becomes fungible. Customers didnโ€™t buy specific IBM machines. They bought into OS/360. The machines were an implementation detail.

Intel understood this by the 1980s, even if implicitly. The 80386โ€™s backward compatibility commitment in 1985 was IBMโ€™s 360 insight applied to microprocessors โ€” the architecture is the product, the silicon is the vehicle. CUDA is the same insight applied to GPU compute. What NVIDIA sold researchers in 2006 wasnโ€™t the G80 card. It was the abstraction: write parallel code in C++, run it on any NVIDIA hardware, trust that the next generation will be faster and compatible.

The pattern is now sixty years old. It has reproduced in every major platform transition. And it keeps working for the same reason it worked in 1964: when you own the layer that developers write to, your customersโ€™ switching costs compound every year they stay.


Thereโ€™s something worth sitting with here. Neither Jensen Huang in 2006 nor Gordon Moore in 1968 could have specified exactly what the payoff would look like. What they shared was a willingness to build infrastructure for a demand they could sense but not yet see โ€” and the discipline to keep investing in it through the long years when it looked like a research project rather than a business.

The question that doesnโ€™t resolve cleanly is whether that kind of patience is a strategy or a personality. And whether, in an industry that now moves faster than the cycles itโ€™s lived through, sixteen-year moats are still the kind that get built.


Which raises the uncomfortable corollary: the same AI tools that CUDA enabled may be what ultimately erodes it.

The attack on CUDAโ€™s moat is now structurally different from anything AMD or Intel could mount before. OpenAIโ€™s Triton compiler lets developers write GPU kernels in Python without touching CUDA at all, and generates optimized machine code that often matches hand-tuned CUDA performance. MLIR โ€” Multi-Level Intermediate Representation, originally from Google โ€” provides a compiler infrastructure that can target any hardware backend from a single codebase. AMDโ€™s ROCm has historically been dismissed as immature; ROCm 7, released this year, delivers meaningfully better inference performance than its predecessors. And perhaps most directly: Claude Code reportedly ported a CUDA codebase to AMDโ€™s ROCm in thirty minutes โ€” work that previously took months of engineering time.

The irony is almost too neat. CUDAโ€™s moat was built on accumulated switching costs: the friction of rewriting code, the library dependencies, the tribal knowledge encoded in a decade of kernel optimizations. AI coding tools are specifically good at exactly that kind of mechanical, high-context translation. The weapon is attacking the wall it was built behind.

That said, itโ€™s worth being careful about the speed of this. Abstraction layers that โ€œshouldโ€ erode moats often take far longer than expected, because the moat isnโ€™t just the code โ€” itโ€™s the ecosystem of tooling, documentation, community knowledge, and hardware-software co-optimization that took eighteen years to compound. Triton and MLIR are real. Theyโ€™re also early. The question isnโ€™t whether the moat is vulnerable; itโ€™s whether it erodes before NVIDIAโ€™s next generation of chips makes it irrelevant to argue about.


As for what comes next โ€” which company is building the IBM 360 of this decade โ€” the honest answer is that itโ€™s too early to call with confidence. But thereโ€™s a candidate worth watching.

Anthropicโ€™s Model Context Protocol, launched in late 2024, has the structural fingerprint of a platform play. MCP is a standard for how AI agents connect to external tools and data sources โ€” a common interface layer, hardware-agnostic (or rather, model-agnostic), that any system can implement. By late 2025 it had been donated to the Linux Foundation, adopted by OpenAI and Google, and was tracking 97 million monthly SDK downloads. There are now over 10,000 MCP servers. It is becoming the way agents talk to the world.

The parallel to OS/360 is imprecise but instructive. What IBM built in 1964 was a standard interface between software and hardware that decoupled what you wrote from what you ran it on. MCP is attempting something similar one abstraction layer higher: decoupling what an agent does from the specific models, tools, and data sources it does it with. If it becomes the standard โ€” the layer that developers write to โ€” then whoever owns or most deeply shapes that standard controls the integration tax of an industry whose applications we canโ€™t fully specify yet.

The counterargument is that open standards, once donated to foundations and broadly adopted, donโ€™t generate the same lock-in as proprietary platforms. OS/360 was IBMโ€™s. CUDA is NVIDIAโ€™s. MCP is now the Linux Foundationโ€™s, with OpenAI and Google as co-stewards. The historical pattern suggests the moat accrues to whoever owns the layer, not whoever invented it.

Which may mean the next great platform play is still being assembled in a room we havenโ€™t seen yet โ€” the way IBMโ€™s System/360 was being architected in a Connecticut motor lodge in 1961, three years before anyone else knew what was coming.

Categories
AI

The Ghost of Edison in the AI Data Center

For over a century, the story of modern electricity has been framed by the “War of the Currents.” Thomas Edison championed Direct Current (DC)โ€”a stable, continuous flow of energyโ€”while Nikola Tesla and George Westinghouse backed Alternating Current (AC), which could be easily stepped up in voltage to travel long distances across the grid.

Tesla won. AC became the lifeblood of the global power grid. But history has a funny way of looping back on itself. Today, as we stand on the precipice of the largest infrastructure build-out in human historyโ€”the artificial intelligence data centerโ€”Edisonโ€™s DC power is making a quiet, monumental comeback.

The catalyst? The sheer, unyielding physics of energy consumption.

The AI boom, driven by massive GPU clusters from companies like NVIDIA, is extraordinarily power-hungry. We are no longer measuring data center power in megawatts; we are measuring it in gigawatts. And when you are dealing with power at that scale, the friction of legacy architecture becomes a multi-billion-dollar bottleneck.

On X Ben Bajarin cited a recent conference discussion by an executive from power management supplier Eaton that highlighted a massive architectural shift happening right now behind the scenes:

“800-volt DC to the rack is probably one of the biggest architectural changes that are starting to be designed into data centers, and a lot of those designs are taking place right now. You know, honestly, when look at Eaton, I think that’s one of the untold stories here, is that DC power is probably one of the biggest transformational things that are going to hit the electrical industry since, quite frankly, AC electricity was around in the Edison days.”

To understand why this is revolutionary, you have to look at how a traditional data center gets its power. Power arrives from the utility grid as medium-voltage AC. It is then stepped down to low-voltage AC, sent to the server floor, converted into DC, stepped down again, and finally fed into the server rack at 54 volts.

Every time power is converted from AC to DC, or stepped down through a transformer, there is a penalty. It generates heat, and it loses energy.

“We estimate that there’s roughly about 5% electrical loss during that transition. If you could just go from DC, directly from the utility feed, all the way through the data center into the rack, that’s 5% efficiency gain that you could get.”

In the abstract, 5% sounds like a rounding error. But scale changes everything. Eaton projects that the upcoming data center build-out to support AI will require somewhere between 50 and 100 gigawatts of power.

“So on 50 gigawatts or 100 gigawatts of power generation that’s needed, that’s 5 gigawatts of power that all of a sudden just appears from the existing infrastructure. And that is really, that is really exciting.”

Five gigawatts is not a rounding error. Five gigawatts is the equivalent output of five standard nuclear reactors. It is enough energy to power millions of homes. And in this new 800-volt DC architecture, those five gigawatts aren’t created by burning more coal, building more solar panels, or splitting more atoms.

They are created purely by the removal of friction. By subtracting the unnecessary steps.

There is a profound philosophical metaphor hidden in this electrical engineering triumph. In our own lives, and in our organizations, we are obsessed with generation. When we face a deficitโ€”a lack of time, a lack of output, a lack of revenueโ€”our default instinct is to generate more. We try to work longer hours, hire more people, or drink more coffee.

But how much of our daily energy is lost to “conversion friction”? How much mental power evaporates when we constantly context-switch between tasks, essentially converting our mental state from AC to DC and back again? How much organizational momentum is lost translating an idea through five different layers of middle management before it reaches the “rack” where the actual work is done?

Often, the most elegant and impactful solution isn’t to generate more power. It is to look at the existing architecture of your life or business, identify the transition points that are bleeding energy as heat, and rewire the system to flow directly to the source.

The invisible architecture that shapes our digital lives is shifting. In the race to build the future of artificial intelligence, the biggest breakthrough wasn’t a new way to create energy, but a century-old method of preserving it.

Categories
AI Software

The Thermodynamics of Thought

For the last two decades, we have lived in the era of zero marginal cost. The defining characteristic of the internet age was that once software was written, distributing it to the billionth user cost virtually the same as distributing it to the first. We grew accustomed to the economics of abundanceโ€”infinite copies, infinite reach, lightweight infrastructure.

But the recent commentary regarding the true nature of Artificial Intelligence forces a jarring mental correction:

“AI is not software riding on old infrastructure. It is a new industrial system that converts energy into intelligence – requiring a capital stack measured in trillions, not billions.”

This distinction is not merely semantic; it is physical.

When we view AI through the lens of traditional SaaS (Software as a Service), we miss the magnitude of the shift. We are looking for an app; what is being built is a refinery. We are witnessing a return to heavy industry, but the commodity being refined isn’t crude oilโ€”it is information, and the byproduct is reasoning.

This requires us to think less in terms of code and more in terms of thermodynamics. In this new industrial system, intelligence is an energy-intensive output. Every token generated, every inference drawn, requires a specific, measurable conversion of electricity into heat and computation. Unlike the static code of a website, an AI model is a furnace. It must be fueled constantly.

This explains the capital stack. We are seeing numbers that seem irrational in the context of venture capitalโ€”trillions, not billions. But if you view a data center not as a server farm, but as a power plant that generates intelligence, the numbers align with historical precedents. We are not funding startups; we are funding the modern equivalent of the electric grid, the transcontinental railroad, or the petrochemical complex.

We are pouring concrete, smelting copper, and manufacturing silicon on a planetary scale. The “cloud” was always a misleading metaphorโ€”it sounded fluffy and ethereal. The reality of the AI transition is heavy, hot, and incredibly expensive.

We are moving from an era where we organized the world’s information (low energy) to an era where we synthesize new reasoning (high energy). We are building a machine that eats electricity and excretes intelligence. That isn’t a software update; that is a new industrial revolution.

Categories
AI Books Nvidia

The Thinking Machine

Over the weekend after Christmas, I started reading Stephen Witt‘s book The Thinking Machine: Jensen Huang, Nvidia, and the World’s Most Coveted Microchip which was published last April.

For some reason, I ignored this book until the end of the year – but wow – was I hooked once I started reading it a few days ago. The book grew out of a New Yorker piece Witt wrote in 2023 titled “How Jensen Huangโ€™s Nvidia Is Powering the A.I. Revolution“.

Witt’s book is obviously about Nvidia and CEO Jensen Huang but it’s also about so much more of what’s happening in the world of AI.

In addition, the last chapter is quite a capstone to the whole book – a delight.

Highly recommended!