Categories
AI Business Technology

The Diffusion of Ordinary Work

A recent O’Reilly Radar piece has stayed with me longer than most: Jeff Ding’s diffusion theory of great-power competition applies just as well to AI adoption, and it suggests that companies chasing the frontier might be optimizing for the wrong thing.

Ding, a political scientist at George Washington University, pushes back on the standard story of technological power — that the country or company which first invents or dominates a glamorous new sector locks in lasting advantage. The historical record says otherwise. General-purpose technologies like steam, electricity, and computing produced durable national advantage not through invention but through diffusion: the slow, unglamorous work of embedding a technology into ordinary productive work across an entire economy. The infrastructure that mattered was never the breakthrough lab. It was the education and training systems that produced large numbers of competent, ordinary engineers who could put the technology to work. Ordinary engineers, in Ding’s framing, matter more than heroic inventors.

The same logic holds inside a company. Frontier models turn over every few months. Organizational know-how compounds.

Palantir makes the abstraction concrete. The company doesn’t train frontier models — it builds the layer underneath them: a live, machine-readable model of how a specific organization actually works, a data integration fabric, and a platform that connects whatever model a customer chooses to real operational decisions. It is deliberately model-agnostic. The value proposition is governance, context, and the accumulation of reusable logic rather than access to the newest weights. Practitioners embed with the customer, learn the domain, and configure the system against the customer’s own data and processes — diffusion as a job description.

Leadership has been unusually blunt about what this implies: frontier labs, they argue, are optimizing for benchmarks while under-delivering on what enterprises actually need. The clearest evidence for the argument is also the most citable one — there have been production cases where an unmodified open-weight model, running inside Palantir’s platform with customer-specific context, outperformed frontier models on the actual task. If true, and it appears to be, the implication is uncomfortable for anyone selling model quality as the whole story: the ground underneath the model — the ontology, the data, the accumulated rules — often determines outcomes more than the model itself.

Electrification is the closest historical analogue. Factories didn’t get more productive the day they installed electric motors. The gains showed up years later, once entire production systems had been redesigned around decentralized power. The lag was organizational, not technical. AI diffusion looks likely to follow the same shape — the bottleneck was never going to be model capability, it was going to be the patient, unglamorous work of redesigning how people actually work.

I don’t know who’s training the ordinary engineers right now — the ones who will spend the next decade doing the diffusion work rather than the invention work. I don’t think anyone’s tracking their names.

Categories
AI

The Quiet Trade-offs of Open Weights

An open letter is circulating this week — Open Weights and American AI Leadership — signed by a broad coalition of companies arguing that downloadable model weights are essential to U.S. competitiveness, diffusion of capability, and even safety. It makes a strong case on access, competition, and sovereignty. It also nods, briefly, to the fact that once weights are released they pass beyond the original developer’s control.

What it doesn’t fully reckon with are two structural realities that follow from that release. Neither is an argument against open weights. Both are simply facts about what openness costs, and what it buys.

Two core limitations

First, control.
Once the weights leave the developer’s servers, the developer can no longer dictate how the model is used. System prompts, refusal training, monitoring, rate limits, rapid safety updates — none of it reaches an independent deployment. Users can strip safeguards, fine-tune for purposes the original team would never sanction, or run the model somewhere it was never meant to go. The letter acknowledges the loss of control. It doesn’t linger on what that means for ongoing safety governance.

Second, learning.
Closed, hosted models draw on a continuous stream of real usage — the queries people actually ask, the reasoning traces that result, the places the model fails or succeeds in the wild. As appropriate that exhaust can be sampled, reviewed, and fed back into improvement. Open weights running independently offer no such path. The developer has no visibility into how the model is being used at scale once it’s out the door. Improvement then falls to slower, thinner channels: community datasets, published evals, distillation from any parallel closed models the lab still runs, internal preference data. The high-volume, real-distribution signal is gone.

These two limitations travel together. The same openness that strips the developer’s control also strips its ability to learn from the model’s actual use.

Sovereignty flips the perspective

A parallel argument has been building around “sovereignty” — an enterprise or government’s ability to own its data, its fine-tuned weights, its compute, its proprietary edge. In this framing, open weights are a path to control, but for the user, not the developer. The organization downloads the model, adapts it inside its own environment — often air-gapped — and keeps whatever capability results private. What the lab surrenders in ongoing control, the institution gains in independence.

But the same move that delivers sovereignty deepens the learning problem. An organization running the model under genuine sovereignty keeps its queries, reasoning traces, and institutional knowledge inside its own walls, by design. None of that returns to the developer. The more high-value users — governments, defense, critical infrastructure, large enterprises — choose sovereign deployments, the thinner the real-world signal available to the labs training the next generation of models. Local fine-tuning can still happen, but that learning stays private. It doesn’t flow back into the shared base model.

What the letter leaves out

The letter is right that closed models aren’t automatically safer, that concentration creates single points of failure, and that transparency invites broader scrutiny. It’s also right that open weights expand access and cut lock-in. Those points hold.

But it treats the developer’s loss of control mainly as a manageable risk that community examination can offset. It celebrates user control and sovereignty without mapping the full exchange: the developer loses both control and its richest usage signal, and that signal thins further as more institutions choose real sovereignty. The information environment models improve in is changed by these choices — not just the distribution of access.

Other distinctions worth naming

  • Update velocity. Closed models patch globally and immediately. Open-weight deployments lag; many users never leave an old version.
  • Customization power. The flip side of lost control is real specialization — downstream users can adapt a model far deeper into a narrow domain than its original developer ever will.
  • Transparency versus opacity. Open weights let outside researchers inspect and red-team a model in ways closed systems don’t allow.
  • Economic structure. Open weights commoditize the base model and push value toward data, fine-tuning, infrastructure, and applications.
  • Privacy at the edge. Running a model fully offline or on private infrastructure is a guarantee hosted services simply can’t match.

A clearer accounting

Open weights aren’t a free lunch. They’re a deliberate trade: the developer gives up ongoing control and the continuous signal of real usage, in exchange for diffusion, customization, outside scrutiny, and user independence. Institutional sovereignty amplifies one side of that trade — it solves the dependency problem for the user while further starving the developer of high-stakes, real-world feedback.

That trade may still be the right one for research progress, economic diffusion, spreading capability beyond a handful of labs, privacy-preserving deployment. But it’s a trade with real, compounding costs. Treating the loss of control as a footnote, and the loss of the learning signal as invisible, leaves an incomplete map.

The letter is right that American leadership will be judged by the strength of the whole ecosystem, not by any single frontier model. An accurate map of that ecosystem has to include what openness and sovereignty actually cost the original developers, in control and in learning both. Only then can we reason clearly about when those costs are worth paying — and what might offset them.

The conversation is better when we name the full set of trade-offs instead of talking around them.

Categories
AI

The Kitchen, Not the Farm

There is a sentence buried in Thinking Machines Lab’s release notes for Inkling, its first proprietary model, that most companies would never let out the door. Describing their own creation, the company states plainly that Inkling is “not the strongest overall model available today, open or closed.”

Read that again. A startup that raised two billion dollars in seed funding at a twelve-billion-dollar valuation, founded by OpenAI’s former CTO and staffed with veterans of the labs currently locked in the most capital-intensive arms race in corporate history, shipped its debut model with an admission of inferiority attached to the label. Not buried in a footnote. Stated in the announcement.

It seems like this was the clearest signal yet that the frontier-capability race may be the wrong game, and that durable value in enterprise AI accrues not to whoever has the smartest model, but to whoever owns the layer where that model gets adapted to a particular customer’s purpose.

As I’ve thought about it, the AI industry seems to be stratifying into three distinct businesses, each with different economics, occupied by a different cast of companies.

It begins with the farm, where the raw ingredients get grown. Then there’s the kitchen, the capital equipment that makes skilled cooking possible at scale. Lastly there’s the restaurant, where somebody who understands a specific customer takes the ingredients, uses the kitchen, and puts a particular dish in front of a particular diner who is paying for a complete meal, not just the flour or the vegetables. Thinking Machines seems to me like a clear example of a company trying to explain which of those businesses it’s actually in. It is not the only one.

The sequence, read backward

Founded in February 2025. Silent for over a year. Then, last October, the company’s first product emerged — and it wasn’t a chatbot, wasn’t an assistant, wasn’t anything a consumer like me would recognize or understand. It was Tinker, a fine-tuning API. Infrastructure for customizing other people’s models, shipped before the company had released a model of its own.

That sequencing is the tell. A company chasing frontier supremacy builds the model first and the tooling around it later, the way the frontier AI labs have all done. Thinking Machines inverted the order. It built the workshop before it built anything to put in the workshop window, which only makes sense if the workshop was always the product.

Inkling, released this month, doesn’t reverse that logic. It completes it. The model is described in the company’s own materials as “an extremely knowledgeable, generalist base that can be extended via fine-tuning” — language that positions the model itself as raw material, not a finished good. It ships with full open weights, day-zero availability on Tinker, and a name chosen, according to the company, to evoke “an idea in its earliest stage, with the potential to grow into something greater.” Even the naming is a thesis statement. Inkling is not meant to be the only thing you use. It’s meant to be the thing you start from.

In the farm-kitchen-restaurant frame, it seems like Thinking Machines is trying to own two levels of the stack at once. Inkling is the farm — grown at real expense, forty-five trillion tokens of training data, frontier-scale compute. Tinker is the kitchen — the induction range and the walk-in fridge, sold as a service to whoever wants to cook. What Thinking Machines has explicitly declined to be, by its own admission, is the restaurant. They are not trying to serve you the best possible dish. They are trying to make sure that whoever does serve you that dish is buying their ingredients and standing at their stove and cooking in their kitchen.

The manifesto that preceded the model

A company doesn’t back into a strategy this coherent by accident. Earlier this month — before Inkling shipped — the lab published a position paper arguing that most AI today is trained in a handful of places and then frozen, a design that by its nature excludes the people the model is meant to serve. Their proposed alternative: AI that is distributed, customizable, and shaped by the people using it, not the lab that built it.

Mira Murati has said the same thing more plainly, and said it a year before Inkling existed, back when Tinker launched. Her framing wasn’t about building the smartest model. It was about making “frontier capabilities much more accessible to all people” — democratization as the mission, not capability supremacy. That is a genuinely different objective function than the one driving her former employer, and it was declared outright, not discovered after the fact to explain a disappointing benchmark result.

Inkling is a 975-billion-parameter mixture-of-experts model trained on forty-five trillion tokens across text, image, audio, and video, with a context window stretching to a million tokens. That is frontier-scale compute expenditure. This isn’t a company that ran out of runway and settled for a smaller ambition. It’s a company that spent frontier-level resources and then declined to spend the final increment chasing benchmark supremacy, presumably because the return on that increment doesn’t show up in the business they’re building.

A second detail: Inkling reportedly uses one-third the tokens of Nemotron 3 Ultra to hit equivalent performance on agentic coding benchmarks. That’s not a capability retreat — that’s a capability choice, optimizing for efficiency and cost-per-task rather than raw benchmark position. And the company is previewing a smaller sibling model alongside Inkling, suggesting a family strategy across sizes rather than a single mid-tier release.

The restaurant next door: Palantir

Thinking Machines isn’t the only company making this bet — and looking at who else is making it shows not everyone is occupying the same layer.

Earlier this month Palantir and Nvidia announced a “Sovereign AI Operating System” — Nvidia’s open Nemotron models, fine-tuned on a customer’s own data, running on Nvidia hardware inside that customer’s own air-gapped network, with Palantir’s Ontology and Foundry software layered on top. CEO Alex Karp pointed out that his enterprise customers don’t want to risk sharing their IP with frontier model providers and asked simply why wouldn’t they control the weights?

It’s tempting to read this as the same argument Thinking Machines is making. It isn’t, quite. What Palantir is selling is the restaurant: the finished, seasoned, plated product — an air-gapped AI system wired into a specific government agency’s or enterprise customer’s actual workflows, with “you control the weights” as the pitch that closes the deal. Palantir isn’t growing wheat. It’s the chef, working with ingredients somebody else grew. Somebody who could be trusted.

Another restaurant: Sierra

Sierra, Bret Taylor and Clay Bavor’s customer-support agent company, makes the same choice even more starkly. Sierra’s own technical writing describes a “constellation of models” architecture: rather than betting on a single LLM, Sierra routes each task inside a customer-service agent to whichever model — from OpenAI, Anthropic, Meta, or elsewhere — handles it best, and explicitly says it invests “in fine-tuned models where off-the-shelf models fail to meet our constraints.” Fine-tuning shows up in Sierra’s stack as one tool among several, alongside retrieval and layered “supervisor” models that catch mistakes before a customer sees them. Sierra has no interest in being a model company or an infrastructure company. It wants to be the restaurant that happens to keep a few specialty ingredients in the walk-in that nobody else stocks, because the dish needs them and they know just how to include them.

Mapping the rest of the stack

The farm-kitchen-restaurant split shows up everywhere the fine-tuning economy has organized itself.

The kitchen-builders — companies selling fine-tuning infrastructure to whoever wants to cook with it, indifferent to what gets made — now form a crowded field: Thinking Machines’ Tinker, Together AI, Fireworks AI, Predibase, OpenPipe, Baseten, Modal, Databricks’ Mosaic stack, and newer entrants like Nebius’s Token Factory and Prime Intellect. None of them care whether you’re building a coding agent, a legal research tool, or a customer-service bot.

The restaurants — companies where fine-tuning is invisible plumbing inside a finished, vertical product — include Palantir and Sierra, and many others. The addressable market for fine-tuning seems to include almost every possible enterprise adopting AI.

What’s seems unusual about Thinking Machines is that it’s trying to be the farm and the kitchen simultaneously while declining, by its own public admission, to be the restaurant. Most companies pick one layer and defend it. Thinking Machines is betting that owning two of the three is the more durable position — grow the flour, own the stove, and let Palantir, Sierra, and a thousand enterprise engineering teams fight over who plates the dish.

The same stack, built by design

As I was thinking about this, I wondered how this relates to the AI activities underway in China. It seems that China’s AI industry maps onto this same three-layer structure with unusual clarity — and one genuine wrinkle the American version doesn’t have.

The farm is crowded and innovating on a different axis than size: DeepSeek, Alibaba’s Qwen, Zhipu AI, Moonshot AI, MiniMax, ByteDance’s Doubao and Seedance. The standout isn’t scale, it’s efficiency — DeepSeek’s V3.2 reportedly uses a novel sparse attention mechanism to nearly match GPT-5 and Gemini 3 on complex reasoning despite far less compute, a different kind of farming: not more wheat, but wheat bred to need less water. Qwen has become the default soil for the rest of the world’s kitchens, generating over 100,000 derivative fine-tunes on Hugging Face. VC’s in Silicon Valley note how frequently their startup companies are building on Qwen.

The kitchen layer has its own SiliconFlow — a Beijing infrastructure startup, backed by Alibaba Cloud, that bills itself as the neutral layer between AI applications and hardware. It solves a problem others never had to: China’s compute runs across fragmented domestic chips, Huawei’s Ascend line chief among them, that don’t share Nvidia’s CUDA ecosystem. SiliconFlow abstracts that fragmentation away — it became the fastest platform serving DeepSeek traffic, and the only large provider running DeepSeek on Ascend chips instead of Nvidia’s. That’s a stove engineered to burn whatever fuel is in the tank that week, a direct product of the U.S. chip export controls rather than any inherent technical edge. Volcano Engine, Alibaba Cloud’s PAI, and Baidu’s Qianfan are versions of the same layer.

The restaurant layer is where China’s picture diverges most from Palantir and Sierra’s venture-funded improvisation: it’s named industrial policy.

Beijing’s “AI+” initiative targets seventy percent sectoral AI penetration by 2027, ninety by 2030 — fine-tuned vertical deployment treated the way past five-year plans treated high-speed rail. The players read like a sector directory: SenseTime for vision and embodied AI, iFlytek for speech in education and government, Baichuan Intelligence for healthcare, 4Paradigm for finance and industry, each fine-tuning a general base into something that only makes sense inside one workflow — a hospital’s diagnostic support tool, a bank’s risk model, an industrial inspection line.

The bet

Every frontier lab is implicitly betting that intelligence is the scarce resource, and that whoever has the most of it wins the enterprise market by default. Thinking Machines, Palantir, Sierra, and many others are all, in their different ways, betting against that premise — that raw intelligence is commoditizing faster than the frontier labs’ spending would suggest, and that the scarce resource has already migrated to whichever layer turns a generalist model into a specific customer’s model.

Thinking Machines is betting the moat moved to the farm-and-kitchen layer. Palantir, Sierra and others are betting it moved further still, to the restaurant, where nobody cares whose flour was used as long as the dish is right. China is betting on all three layers at once, with the state underwriting the bet directly.

It is a curious thing for me to watch companies with this much money and this much talent choose not to fight for the title of smartest model in the room. It is also a curious thing to watch them explain why, in public, in the first paragraph of an announcement.

But I think I’m beginning to understand.

Categories
AI Business

The Reverse Information Paradox We’ve Always Had

Satya Nadella wrote recently about what he calls the Reverse Information Paradox: enterprises pay for AI intelligence twice. Once in money. Again in the proprietary knowledge they surrender through every prompt, correction, and evaluation. The better they use the model, the more of their own institutional understanding leaks into someone else’s system. The vendor ends up knowing more about the buyer’s business than the buyer knows about what the vendor retained.

Replace “model” with “employee” (or “consultant”) and the paradox is not new at all.

You pay for a person once with salary. You pay again with something harder to price: the context, relationships, and judgment they must absorb to become useful to you. The better they perform, the deeper the immersion, the more of your particular way of doing things moves into their head. Every correction and late-night conversation is another trace of institutional memory changing hands. When they leave, some of that memory leaves with them. Not always through theft. Usually just through the ordinary residue of good work.

The visible cost is salary; the invisible cost is the slow transfer of what makes you distinctive. High performers get more access precisely because they’re high performers, which means the leakage accelerates exactly when you can least afford it. The exhaust is just harder to see with people than with tokens — it moves through conversation and mental models instead of logs.

The analogy has a limit, and the limit matters. Employees bring knowledge in, not just absorb it. They have judgment and relationships a model doesn’t. Models are purely absorptive, and once something is inside them, it’s infinitely reproducible — a person can only be in one place, working for one employer, at a time. We’ve had a few hundred years to build tools for the human version of this problem: contracts, culture, non-competes. The model equivalent is still being invented in real time, which is exactly why Nadella felt the need to name it.

Apple’s recent legal action against former employees who joined OpenAI is this pattern in its sharpest form. Whatever the specifics, the shape is familiar: people who spent years inside one of the most sophisticated organizations in the world, carrying out knowledge that never appeared on any balance sheet and was hard to contain. No one fully anticipates what a mind absorbs simply by being in the room long enough.

That’s the real difference between the silicon case and the human one. You can try to take action to wall off knowledge flowing to a model. You cannot wall off what someone has learned to notice.

Categories
AI Photography

The Price of the Cold

Two men are standing close to a brick wall trying not to talk, because talking wastes what little warmth is left in a body that has been outside too long. One of them has a camera — Jerry Schatzberg, a fashion photographer. His hands are jammed half into his coat pockets between shots. The other man has his collar up around his ears and a scarf wound twice, black and white, and he is not moving much, because moving costs heat, and heat is the one thing neither of them has enough of. Schatzberg raises the camera. His fingers, by this point, are not entirely his own. When he presses the shutter there is a tremor in it he did not order and cannot undo.

The picture comes out smeared at the edges. Bob Dylan’s face, in the frame, is dissolving slightly into the gray behind him, like a man photographed through a windshield in the rain. It is, by any studio standard, a bad photograph. Schatzberg knows it’s a bad photograph. He has made a career out of not taking bad photographs.

And it became the cover of Blonde on Blonde, which is the best rock album ever recorded, and in nearly sixty years nobody has managed to improve on it by reshooting it clean. The blur isn’t a decision. It’s a symptom — of two men standing in the cold too long, of a photographer choosing, afterward, to keep the evidence of his own discomfort instead of erasing it.

There’s a difference between an accident and serendipity that I don’t think gets said out loud enough, and it matters more than it used to. An accident is the cold — involuntary, uninvited, spent before you know if it was worth spending. Schatzberg didn’t choose to shiver. His hands moved because his body was doing what bodies do at a certain temperature, and the shutter caught what his hands actually did, not what he meant to do. Serendipity is what happens next: a verdict, rendered after the fact, that the wreckage of an intention was better than the intention itself. The accident is what makes the verdict possible. Without the cold, there’s nothing to render a verdict on.

I’ve been sitting with a large language model most days for the better part of a year now, watching it write, asking it to try again, watching it try again in a way that is never quite the same and never quite different enough to matter. Somewhere upstream of me there is a number called temperature, and I will never see it. Somebody else did, once, in a meeting, and decided that the word for controlled, pre-approved, refundable randomness should be temperature — the same word for the thing that made Schatzberg’s hands shake, the same word for the actual physical stakes of standing outside too long in January without enough coat — and then set it, and moved on, and nobody in that meeting laughed, because nobody in the room had ever been cold in a way that mattered to the work.

Picture the room instead. It is climate-controlled to sixty-eight degrees, humidity held flat, year-round, by a building management system nobody thinks about until it fails. Somewhere in it, the hardware is generating your next five versions of a photograph like the one on Blonde on Blonde. Nobody in that room is going to lose feeling in their fingers today. Nobody’s collar is up. I don’t know his name — nobody outside the building does — but somebody like him tuned the sampling distribution and went home at six. That’s the guy in the good suit. He built the weather. He never once stood in it.

The small model inherits conclusions. It never inherits the cold. Whatever accidents shaped the teacher model’s own training — whatever costly friction produced the insight in the first place — the student model gets none of that weather. It gets the photograph, cropped and sharpened, with the blur removed because somebody along the way decided the blur was noise instead of signal — the way Schatzberg, a lesser photographer, might have reshot Dylan clean and thrown the bad one away. It is heir to a serendipity it never earned, because it was never present for the accident that made the serendipity possible. It is, in the most literal sense the industry means by the word, cheap.

I keep coming back to the fact that nobody at the API layer is shivering. That’s not a complaint, exactly. It’s just an observation about where the cost went. Somewhere in the training data, some human being was cold, or scared, or holding a fish that was starting to smell, or standing on a stepladder with ten minutes before the traffic came back, and that person paid a real price for a result they couldn’t yet know was good. The model downstream of all that gets the result without the price.

Two rooms, then. In one of them it is January in New York and a man’s fingers have stopped entirely obeying him. In the other it is sixty-eight degrees, always, on a Tuesday and on a Sunday and at three in the morning, and the machines are making you nine more versions of that same blur. Sixty-eight degrees. A number, upstream, that you will never see.

Categories
AI AI: Large Language Models Apple

The Slipstream Strategy

Apple had a problem no amount of money could solve. An iPhone can’t draw the power or shed the heat of a data center, so ten different tasks can’t mean ten different models fighting for the same sliver of RAM. Apple’s answer was to freeze one small, efficient base model into the device and then swap tiny adapters in and out of it in milliseconds — a summarization adapter for your texts, a Siri adapter for on-screen actions, and a handoff to Private Cloud Compute for anything heavier. The phone behaves like it’s running many models. It’s running one model wearing many hats.

That architecture — a frozen base plus swappable adapters — is quietly becoming the default way serious AI companies build, and it’s worth understanding why, because it inverts the assumption most people still carry into this industry.

The assumption is that winning means owning a frontier model. Sierra co-founder Clay Bavor pushed back on that on a recent 20VC episode: pouring capital into your own pre-training, he argued, tends to leave you holding a highly perishable bag of floating-point numbers. Open-weight models improve fast enough that yesterday’s frontier is next quarter’s commodity. The companies playing this well aren’t racing to out-spend the labs. They’re slipstreaming behind them — taking the free, state-of-the-art engine and putting all their effort into what sits on top of it.

What sits on top is LoRA — low-rank adaptation. The old failure mode was catastrophic forgetting: fine-tune a model hard enough on your own data and it forgets how to reason generally. LoRA sidesteps this by leaving the base model untouched and training a small set of additional parameters alongside it — a thin layer of expertise bolted onto a frozen foundation. You get real domain depth without touching the thing that makes the model work at all.

The business logic that follows from this is the actual point, and it’s simpler than it looks:

You stop being hostage to any one model provider — if a better open-weight model ships next month, you port your adapter, not your whole product. You can serve hundreds of differently-customized clients off one base model on one piece of hardware, instead of running a separate giant model per customer. You can ship a fix in an afternoon, because an adapter is a few hundred megabytes, not a training run. And in regulated industries, your proprietary data can train an adapter that never leaves your own infrastructure.

None of this is really a story about model architecture. It’s a story about where the moat moved. For a while the moat was raw capability — whoever had the best model won. Apple and Sierra are betting the moat is now somewhere else entirely: in how tightly you can weave a commodity intelligence into a specific workflow, a specific dataset, a specific customer relationship. The engine is free. The adapter is the business.

Categories
AI

Context Rot

Here is a small, possibly embarrassing confession: I have never, not once, gone looking for the best AI model.

I have a model. It lives in a browser tab — Safari, usually, on whichever device is nearest, occasionally Chrome if I happen to be at the desktop. It does what I need — drafts an email, untangles a sentence, tells me what a Norwegian emigration record from 1856 probably says — and then I close the tab and go on a walk.

Somewhere out there, presumably, a much smarter, much more expensive machine is doing something extraordinary with protein folding or hedge fund arbitrage or the outer edges of mathematics I will never visit. I have made my peace with never meeting it.

This did not used to feel like a confession. For a while there — a year, eighteen months — it felt like the central drama of the whole industry: which model was “best,” who had it, who had lost it, whether some lab’s quarterly earnings call would reveal that the frontier had quietly moved sixty miles down the road while everyone was looking the other way. Benchmarks were released like box scores. People argued about them the way people argue about batting averages, with the same weird intensity, the same conviction that a two-point difference in some abstract reasoning test settled something important about the future.

And then, at some point I can’t quite date — it crept up, the way these things do — I noticed I had stopped caring.

Not because the frontier stopped moving. It didn’t. It’s still moving, arguably faster than ever, in ways that occasionally show up in the news with all the drama of a soap opera (a delayed launch, a researcher poached, a stock down five percent in an afternoon, always something).

I stopped caring because none of it touched me. My model — whatever it was, this week — had long since crossed some invisible threshold past which more didn’t register as more. It was already better than I needed. It has been better than I needed for a while now. I suspect I am not unusual in this. I suspect most people, doing most things, most days, are operating comfortably inside a capability surplus so large they’ve stopped noticing it’s there, the way you stop noticing a room is warm.

If the top of the model isn’t for people like me — and it increasingly isn’t — then who, or what, is it actually for? I went looking for one piece of the answer and found, instead, a metaphor.

It’s called “context rot.” I have to admit, before I go further, that I’m not sure I’ve ever felt it myself — which, on reflection, is its own small piece of evidence. My sessions close in minutes, not hours. I ask, it answers, I leave. Whatever happens to a model over the fourth or fifth hour of sustained, dependent work is a country I simply don’t visit.

But other people do, increasingly — entire teams do, for entire projects — and what they’re finding out there is worth understanding, even secondhand. It describes something that happens to AI models when they’re asked to work for a long time on something complicated — not five minutes, but five hours; not one question, but a hundred small decisions stacked on top of each other, each one depending on the last.

You’d think the limiting factor would be room. Models have a “context window” — a stated capacity, like a gas tank, measured in tokens, and for a while the marketing numbers on these were the whole story: two million tokens! A library! And you’d think, as with a gas tank, that the thing runs fine until it’s empty and then it stops.

That is not, it turns out, what happens. What happens is closer to what happens to your desk.

You know the desk. Everyone has the desk. It starts the morning clean — an aspirational, almost insulting cleanliness — and by four in the afternoon it is a geological record of the day: three coffee cups, a stack of things you meant to file, a Post-it with a phone number you no longer need, the good pen buried under a printout of something you already dealt with an hour ago. The desk is not full. There is, technically, room. You could clear a space if you tried. But you don’t try, because functionally, cognitively, the desk has stopped being usable long before it ran out of surface area. You start looking for the stapler and forget what you were stapling. This — and I did not make this term up, I want to be clear, though I wish I had — is context rot. The window hasn’t run out. The signal has just drowned in its own debris.

Researchers watching this happen to long-running AI agents have found something almost cruelly elegant about how it fails: it doesn’t fail gradually, the way you’d expect a desk to get gradually messier. Errors compound. A task that takes twice as long doesn’t get twice as likely to go wrong — the failure rate roughly quadruples. Two mistakes early in a long chain of dependent steps don’t add up to a slightly worse outcome. They multiply into something close to total collapse, four hours in, for reasons that trace back to a single bad assumption made in the first twenty minutes and never revisited.

Here is where the frontier comes back in — not as the whole answer, but as a piece of one.

It is not that frontier models are smarter in the way a benchmark measures smart — better at a single hard math problem, a cleverer turn of reasoning. Plenty of models can do that now; the “good enough” tier has crept remarkably high.

It’s that frontier models are apparently, marginally, meaningfully better at not rotting. At keeping the desk usable at hour six. At knowing which of the forty things on the desk actually still matters and which is a coffee cup that should have been thrown out an hour ago. This is a genuinely different kind of intelligence than the one benchmarks were built to measure, and it is almost invisible from the outside — you don’t see it in a single exchange, you see it only in the difference between a project that holds together over three days and one that quietly, subtly, stops making sense somewhere around Tuesday afternoon and nobody notices until Thursday.

If that’s true — if the frontier’s real edge is durability rather than raw cleverness — you’d expect to see it show up in how the labs actually deploy their own models: saving the sharpest tools for the tasks that need to survive the longest.

I went looking for a real-world example and found one closer to home than I expected: Anthropic’s own Slack tool, the one where you tag the AI into a channel the way you’d tag a coworker, and it works alongside a whole team over days, learning the channel as it goes. It runs on a serious, capable, thoroughly frontier model — but not, it turns out, on the company’s very best one. That one is held back, reserved for a smaller and stranger set of problems nobody has solved before at all. I sat with that for a while. The tool built to survive a whole team’s whole week, in public, under the most sustained pressure any of their products face, wasn’t handed the sharpest blade in the drawer. It was handed the second-sharpest — which was apparently, entirely, enough. Which tells you something about where the two kinds of intelligence actually diverge: the merely-very-good model handles the desk staying clean for a week, in public, in front of a whole team, where one bad assumption made Monday and never revisited would be visible to everyone by Thursday. The truly new capability is being held in reserve for something else altogether.

I don’t have a tidy place to land this, and I’m suspicious of anyone who does. But here’s the closest I can get.

Imagine a three-Michelin-star chef — the kind of person who has spent thirty years learning to coax something transcendent out of a single scallop, who can tell you, by smell, that a stock has forty more minutes in it — standing at your stove on a Tuesday night making you a grilled cheese sandwich. It will, I promise you, be a very good grilled cheese sandwich. The bread will be evenly golden. The cheese will have reached some ideal, fully-considered state of melt. But almost none of what makes that chef extraordinary is actually being used to make it — none of the thirty years spent learning to hold forty things in mind at once without losing track of any of them, the exact skill, it occurs to me, that keeps a long, complicated project from quietly falling apart on day three. The technique is idling. The thirty years are in the room, present, available, and almost entirely beside the point, because a grilled cheese sandwich was never the place where thirty years shows up. It shows up somewhere else — in a dish you will never order, on a night you weren’t there.

What you got instead, on your ordinary Tuesday, was simply more than enough.

Categories
AI Consulting

The Judgment Layer

An analyst’s note about the CEO of one of the largest consulting companies making comments at an investor conference includes a line that deserves more attention than it got: “token volume used on a project isn’t a proxy for AI maturity.”

Translation — clients are burning money on frontier models for problems that don’t need frontier models, and they’re not getting the outcomes they expected.

This firm’s CEO offered this as a business opportunity. I read it as a confession.

The old consulting model was simple: client has a technology problem, firm deploys humans to solve it. Billing followed effort. The new problem is different in kind — clients have an AI strategy problem. They know they’re supposed to be using AI. They’ve heard the word “frontier.” They’re spending accordingly. They just don’t know why, and the outcomes are showing it.

So the CEO is right that there’s an opportunity here. The value proposition shifts from implementation to judgment — not deploying AI, but knowing when not to deploy the expensive one. Matching capability to problem. Being trusted enough to tell a client that their $50M frontier model contract is solving a $500K problem.

Here’s the irony that the comment skates past: that advice is structurally difficult for a large consultancy to give.

The business model that built consulting firms was billing for doing. The more you deploy, the more you bill. Helping a client spend less, or choose the cheaper model, or run a narrower project, is genuinely good advice that the incentive structure actively works against. You don’t grow a $70 billion professional services firm by talking clients out of scope.

The judgment layer, if it becomes the real value, requires something closer to a doctor’s relationship with a patient than a contractor’s relationship with a client. Doctors get paid whether they prescribe or not. The value of the visit is the diagnosis — including the diagnosis that says you don’t need the expensive intervention. Consultants, historically, get paid to prescribe, and paid more when the prescription is larger.

There’s a reason we trust doctors with that asymmetry and not contractors. Licensing, malpractice, professional norms built over centuries — all of it exists to align the incentive. Consulting has none of that infrastructure. What it has instead is reputation, which is slower-acting and easier to game.

Whether the large firms can actually make the shift — rather than just reframe the same billable-hours model in the language of AI optimization — is the real question the market is wrestling with. The CEO’s comment is genuinely perceptive about where client value lies. It’s less clear that consulting firms are currently built to capture it honestly.