Categories
AI Learning Meta

Trajectories: how Meta plans to make Muse smarter by watching it work

Buried in the data policy section of Meta’s long post on how it built safety into Muse is a sentence that isn’t about safety at all:

“Inference data, the back and forth conversations between you and your Muse and the tool calls and subagent handoffs that result (‘trajectories’) are useful data for training new checkpoints of the LLM model at the core.”

Two things are worth noticing. First, the technique described isn’t new — training on agent rollouts is standard practice across the field. Second, the company describing it is Meta, in plain language, in a public post, about a product aimed at billions of consumers. The labs usually discuss this stuff in papers about coding agents. Meta just told its future user base: your agent’s work product is our training data. The candor is the story, not the technique.

A trajectory isn’t a chat log. It’s the complete record of an agent doing a job: what you asked, what it tried, which tools it called, where it went wrong, how it recovered, which subagents it spawned, and whether the thing actually got done. Every time you let Muse book the flight, triage the inbox, or research the supplier, you’re generating one.

From text to behavior

The technique matters anyway, because the diet that AI trains on is changing. Pretraining was about text — the whole internet, more or less. Post-training was about preferences — which answer humans liked better. Trajectories are the third course: demonstrations of competent behavior, in full, mistakes included.

There’s a reason for the shift. Text teaches a model what the world looks like. Preferences teach it what people want. But neither teaches it how to do a 40-step task without wandering off, recovering from a dead end, or knowing when to ask for help. That only exists in records of agents actually doing things. And until recently, almost nobody had those records at scale — because almost nobody had agents doing real work at scale.

The demonstrated instance — and what it doesn’t prove

Meta’s concrete example is Muse Spark 1.2, co-trained with Muse Code: model and harness trained together on rejection-sampled harness trajectories — run the agent many times, keep the runs that succeeded, train on those — with recipe-level tuning for goals, context compaction, and subagents. The model isn’t learning to predict text; it’s learning to behave inside a specific set of tools.

This is the end of the “base model plus clever prompting” era. The artifact is the bundle — model and harness, co-designed. A model trained on trajectories from one harness will be genuinely better inside that harness than a smarter general model dropped into it cold. Meta is saying this out loud; OpenAI and Anthropic are doing the same thing more quietly.

But notice the domain: coding. And coding is exactly where trajectories are cheapest to manufacture — verifiable unit tests, sandboxed repos, SWE-bench-style tasks. Nothing about the Muse Code result requires a single consumer or a single inbox. So the one demonstrated instance of Meta’s trajectory training sits squarely in the category where Meta’s distribution advantage matters least. Meta hasn’t shown its hand on the category that actually matters.

The other category is the personal one, and there the evidence is thinner. What exists is a stated intent, not a published result. The data policy says personal Muse trajectories “are useful data for training new checkpoints.” The product is designed to generate them: Meta’s own design example has Muse monitoring school emails, adding dates to a family calendar, filling a supply cart, finding a sale sweatshirt, booking dinner, and catching a sports tryout deadline hours before it closed. That is what an unverifiable-domain trajectory looks like — a morning of small judgments no unit test could grade.

No training run on that data has been published. No benchmark, no “Muse got X% better at inbox triage after training on Y million user trajectories.” So the sharpest version of the argument — that the real moat is the data nobody else can fake — should be labeled for what it is: a prediction, not an observed fact. It’s a prediction with a mechanism, though: these are judgments that can’t be synthesized, in the one distribution channel that reaches the people making them.

The flywheel — and its limits

With that caveat on the table: trajectories get better with scale, and Meta has scale like nobody else: billions of users across its apps, and now an agent — Muse — sitting inside them. Every user interaction is a potential training trajectory. Better trajectories train a better model; a better model makes a better agent; a better agent attracts more users. Meta states the bargain plainly: “every Muse user gets a better personal agent as we all collectively use the product and help the model understand the intricacies of human life.”

But “most users = most trajectories = structural advantage” needs its counter-case, because a lot of the highest-value trajectory data right now doesn’t come from consumers at all. It comes from sandboxes, the same kind that produced Muse Code. Synthetic and simulated trajectories sidestep the need for billions of users entirely — Anthropic and OpenAI are getting rich trajectory data from developers running Claude Code and Codex against real repos, no social-app distribution required.

The honest version of the moat argument is narrower, and more interesting. Synthetic trajectories work brilliantly where success is verifiable — code either passes the tests or it doesn’t. They work poorly where success is a matter of judgment: triaging an inbox, planning a trip around someone’s actual preferences, knowing which email deserves a reply. There is no unit test for a life well managed. And those unverifiable, deeply personal tasks are exactly what Meta means by “personal superintelligence” — and exactly where its distribution gives it trajectories nobody else can synthesize. The moat isn’t “most data.” It’s “the data nobody else can fake.”

The price of the flywheel

There’s a wrinkle, and Meta knows it. The flywheel runs on your data — your emails, your calendar, the messy reality of your life, which is exactly what makes the trajectories valuable. Meta’s answer is sanitization (“trajectories are sanitized to remove key personally identifiable information”), an opt-out switch, no sharing with ad systems, and a forthcoming “Confidential VM” that would cryptographically prevent even Meta from seeing your data.

The tension is fundamental, and it’s the sharpest part of the whole picture: the product gets smarter by watching you, and it earns the right to watch you by being trustworthy. Those two imperatives pull in opposite directions, and no amount of engineering fully resolves it — the Confidential VM, if it ever ships as described, would resolve it by breaking the flywheel, since trajectories Meta can’t see are trajectories Meta can’t train on. The opt-out rate will be the market’s verdict on the deal Meta is offering.

Experience is the missing piece

But the deepest reason trajectories matter has nothing to do with Meta’s strategy. It’s about what intelligence actually is.

A model trained only on text knows the world the way a brilliant student knows it from books. A model trained on trajectories knows it the way a practitioner does — from doing the thing, failing at it, and adjusting. The trajectory is the closest thing AI has to experience. And an agent that records its experience, keeps what worked, and folds it back into itself is doing something that rhymes with learning.

This is why I keep coming back to continual learning as the critical missing piece in AI. The models are frozen at training time; everything they “learn” afterward lives in context windows and memory files, fragile and local. Trajectories are the bridge: today’s version of the loop is slow and centralized (collect trajectories, train a new checkpoint, ship it), but the direction is obvious. The end state is an agent that learns continuously from its own experience — from your experience with it — the way people do.

Meta’s bet is that the path to personal superintelligence runs through watching agents work, at planetary scale, and distilling what works back into the model. No result yet proves the bet pays off — the personal trajectories are still a hypothesis, not a track record. But it’s an unglamorous hypothesis, no new scaling law, just better data about doing things, and unglamorous bets about data have a good track record in this field. The internet made the last generation of models. Trajectories might make the next one — if Meta can show, and not just say, that the data nobody else can fake is data that actually teaches.

Categories
AI

The Kitchen, Not the Farm

There is a sentence buried in Thinking Machines Lab’s release notes for Inkling, its first proprietary model, that most companies would never let out the door. Describing their own creation, the company states plainly that Inkling is “not the strongest overall model available today, open or closed.”

Read that again. A startup that raised two billion dollars in seed funding at a twelve-billion-dollar valuation, founded by OpenAI’s former CTO and staffed with veterans of the labs currently locked in the most capital-intensive arms race in corporate history, shipped its debut model with an admission of inferiority attached to the label. Not buried in a footnote. Stated in the announcement.

It seems like this was the clearest signal yet that the frontier-capability race may be the wrong game, and that durable value in enterprise AI accrues not to whoever has the smartest model, but to whoever owns the layer where that model gets adapted to a particular customer’s purpose.

As I’ve thought about it, the AI industry seems to be stratifying into three distinct businesses, each with different economics, occupied by a different cast of companies.

It begins with the farm, where the raw ingredients get grown. Then there’s the kitchen, the capital equipment that makes skilled cooking possible at scale. Lastly there’s the restaurant, where somebody who understands a specific customer takes the ingredients, uses the kitchen, and puts a particular dish in front of a particular diner who is paying for a complete meal, not just the flour or the vegetables. Thinking Machines seems to me like a clear example of a company trying to explain which of those businesses it’s actually in. It is not the only one.

The sequence, read backward

Founded in February 2025. Silent for over a year. Then, last October, the company’s first product emerged — and it wasn’t a chatbot, wasn’t an assistant, wasn’t anything a consumer like me would recognize or understand. It was Tinker, a fine-tuning API. Infrastructure for customizing other people’s models, shipped before the company had released a model of its own.

That sequencing is the tell. A company chasing frontier supremacy builds the model first and the tooling around it later, the way the frontier AI labs have all done. Thinking Machines inverted the order. It built the workshop before it built anything to put in the workshop window, which only makes sense if the workshop was always the product.

Inkling, released this month, doesn’t reverse that logic. It completes it. The model is described in the company’s own materials as “an extremely knowledgeable, generalist base that can be extended via fine-tuning” — language that positions the model itself as raw material, not a finished good. It ships with full open weights, day-zero availability on Tinker, and a name chosen, according to the company, to evoke “an idea in its earliest stage, with the potential to grow into something greater.” Even the naming is a thesis statement. Inkling is not meant to be the only thing you use. It’s meant to be the thing you start from.

In the farm-kitchen-restaurant frame, it seems like Thinking Machines is trying to own two levels of the stack at once. Inkling is the farm — grown at real expense, forty-five trillion tokens of training data, frontier-scale compute. Tinker is the kitchen — the induction range and the walk-in fridge, sold as a service to whoever wants to cook. What Thinking Machines has explicitly declined to be, by its own admission, is the restaurant. They are not trying to serve you the best possible dish. They are trying to make sure that whoever does serve you that dish is buying their ingredients and standing at their stove and cooking in their kitchen.

The manifesto that preceded the model

A company doesn’t back into a strategy this coherent by accident. Earlier this month — before Inkling shipped — the lab published a position paper arguing that most AI today is trained in a handful of places and then frozen, a design that by its nature excludes the people the model is meant to serve. Their proposed alternative: AI that is distributed, customizable, and shaped by the people using it, not the lab that built it.

Mira Murati has said the same thing more plainly, and said it a year before Inkling existed, back when Tinker launched. Her framing wasn’t about building the smartest model. It was about making “frontier capabilities much more accessible to all people” — democratization as the mission, not capability supremacy. That is a genuinely different objective function than the one driving her former employer, and it was declared outright, not discovered after the fact to explain a disappointing benchmark result.

Inkling is a 975-billion-parameter mixture-of-experts model trained on forty-five trillion tokens across text, image, audio, and video, with a context window stretching to a million tokens. That is frontier-scale compute expenditure. This isn’t a company that ran out of runway and settled for a smaller ambition. It’s a company that spent frontier-level resources and then declined to spend the final increment chasing benchmark supremacy, presumably because the return on that increment doesn’t show up in the business they’re building.

A second detail: Inkling reportedly uses one-third the tokens of Nemotron 3 Ultra to hit equivalent performance on agentic coding benchmarks. That’s not a capability retreat — that’s a capability choice, optimizing for efficiency and cost-per-task rather than raw benchmark position. And the company is previewing a smaller sibling model alongside Inkling, suggesting a family strategy across sizes rather than a single mid-tier release.

The restaurant next door: Palantir

Thinking Machines isn’t the only company making this bet — and looking at who else is making it shows not everyone is occupying the same layer.

Earlier this month Palantir and Nvidia announced a “Sovereign AI Operating System” — Nvidia’s open Nemotron models, fine-tuned on a customer’s own data, running on Nvidia hardware inside that customer’s own air-gapped network, with Palantir’s Ontology and Foundry software layered on top. CEO Alex Karp pointed out that his enterprise customers don’t want to risk sharing their IP with frontier model providers and asked simply why wouldn’t they control the weights?

It’s tempting to read this as the same argument Thinking Machines is making. It isn’t, quite. What Palantir is selling is the restaurant: the finished, seasoned, plated product — an air-gapped AI system wired into a specific government agency’s or enterprise customer’s actual workflows, with “you control the weights” as the pitch that closes the deal. Palantir isn’t growing wheat. It’s the chef, working with ingredients somebody else grew. Somebody who could be trusted.

Another restaurant: Sierra

Sierra, Bret Taylor and Clay Bavor’s customer-support agent company, makes the same choice even more starkly. Sierra’s own technical writing describes a “constellation of models” architecture: rather than betting on a single LLM, Sierra routes each task inside a customer-service agent to whichever model — from OpenAI, Anthropic, Meta, or elsewhere — handles it best, and explicitly says it invests “in fine-tuned models where off-the-shelf models fail to meet our constraints.” Fine-tuning shows up in Sierra’s stack as one tool among several, alongside retrieval and layered “supervisor” models that catch mistakes before a customer sees them. Sierra has no interest in being a model company or an infrastructure company. It wants to be the restaurant that happens to keep a few specialty ingredients in the walk-in that nobody else stocks, because the dish needs them and they know just how to include them.

Mapping the rest of the stack

The farm-kitchen-restaurant split shows up everywhere the fine-tuning economy has organized itself.

The kitchen-builders — companies selling fine-tuning infrastructure to whoever wants to cook with it, indifferent to what gets made — now form a crowded field: Thinking Machines’ Tinker, Together AI, Fireworks AI, Predibase, OpenPipe, Baseten, Modal, Databricks’ Mosaic stack, and newer entrants like Nebius’s Token Factory and Prime Intellect. None of them care whether you’re building a coding agent, a legal research tool, or a customer-service bot.

The restaurants — companies where fine-tuning is invisible plumbing inside a finished, vertical product — include Palantir and Sierra, and many others. The addressable market for fine-tuning seems to include almost every possible enterprise adopting AI.

What’s seems unusual about Thinking Machines is that it’s trying to be the farm and the kitchen simultaneously while declining, by its own public admission, to be the restaurant. Most companies pick one layer and defend it. Thinking Machines is betting that owning two of the three is the more durable position — grow the flour, own the stove, and let Palantir, Sierra, and a thousand enterprise engineering teams fight over who plates the dish.

The same stack, built by design

As I was thinking about this, I wondered how this relates to the AI activities underway in China. It seems that China’s AI industry maps onto this same three-layer structure with unusual clarity — and one genuine wrinkle the American version doesn’t have.

The farm is crowded and innovating on a different axis than size: DeepSeek, Alibaba’s Qwen, Zhipu AI, Moonshot AI, MiniMax, ByteDance’s Doubao and Seedance. The standout isn’t scale, it’s efficiency — DeepSeek’s V3.2 reportedly uses a novel sparse attention mechanism to nearly match GPT-5 and Gemini 3 on complex reasoning despite far less compute, a different kind of farming: not more wheat, but wheat bred to need less water. Qwen has become the default soil for the rest of the world’s kitchens, generating over 100,000 derivative fine-tunes on Hugging Face. VC’s in Silicon Valley note how frequently their startup companies are building on Qwen.

The kitchen layer has its own SiliconFlow — a Beijing infrastructure startup, backed by Alibaba Cloud, that bills itself as the neutral layer between AI applications and hardware. It solves a problem others never had to: China’s compute runs across fragmented domestic chips, Huawei’s Ascend line chief among them, that don’t share Nvidia’s CUDA ecosystem. SiliconFlow abstracts that fragmentation away — it became the fastest platform serving DeepSeek traffic, and the only large provider running DeepSeek on Ascend chips instead of Nvidia’s. That’s a stove engineered to burn whatever fuel is in the tank that week, a direct product of the U.S. chip export controls rather than any inherent technical edge. Volcano Engine, Alibaba Cloud’s PAI, and Baidu’s Qianfan are versions of the same layer.

The restaurant layer is where China’s picture diverges most from Palantir and Sierra’s venture-funded improvisation: it’s named industrial policy.

Beijing’s “AI+” initiative targets seventy percent sectoral AI penetration by 2027, ninety by 2030 — fine-tuned vertical deployment treated the way past five-year plans treated high-speed rail. The players read like a sector directory: SenseTime for vision and embodied AI, iFlytek for speech in education and government, Baichuan Intelligence for healthcare, 4Paradigm for finance and industry, each fine-tuning a general base into something that only makes sense inside one workflow — a hospital’s diagnostic support tool, a bank’s risk model, an industrial inspection line.

The bet

Every frontier lab is implicitly betting that intelligence is the scarce resource, and that whoever has the most of it wins the enterprise market by default. Thinking Machines, Palantir, Sierra, and many others are all, in their different ways, betting against that premise — that raw intelligence is commoditizing faster than the frontier labs’ spending would suggest, and that the scarce resource has already migrated to whichever layer turns a generalist model into a specific customer’s model.

Thinking Machines is betting the moat moved to the farm-and-kitchen layer. Palantir, Sierra and others are betting it moved further still, to the restaurant, where nobody cares whose flour was used as long as the dish is right. China is betting on all three layers at once, with the state underwriting the bet directly.

It is a curious thing for me to watch companies with this much money and this much talent choose not to fight for the title of smartest model in the room. It is also a curious thing to watch them explain why, in public, in the first paragraph of an announcement.

But I think I’m beginning to understand.

Categories
AI AI: Large Language Models Apple

The Slipstream Strategy

Apple had a problem no amount of money could solve. An iPhone can’t draw the power or shed the heat of a data center, so ten different tasks can’t mean ten different models fighting for the same sliver of RAM. Apple’s answer was to freeze one small, efficient base model into the device and then swap tiny adapters in and out of it in milliseconds — a summarization adapter for your texts, a Siri adapter for on-screen actions, and a handoff to Private Cloud Compute for anything heavier. The phone behaves like it’s running many models. It’s running one model wearing many hats.

That architecture — a frozen base plus swappable adapters — is quietly becoming the default way serious AI companies build, and it’s worth understanding why, because it inverts the assumption most people still carry into this industry.

The assumption is that winning means owning a frontier model. Sierra co-founder Clay Bavor pushed back on that on a recent 20VC episode: pouring capital into your own pre-training, he argued, tends to leave you holding a highly perishable bag of floating-point numbers. Open-weight models improve fast enough that yesterday’s frontier is next quarter’s commodity. The companies playing this well aren’t racing to out-spend the labs. They’re slipstreaming behind them — taking the free, state-of-the-art engine and putting all their effort into what sits on top of it.

What sits on top is LoRA — low-rank adaptation. The old failure mode was catastrophic forgetting: fine-tune a model hard enough on your own data and it forgets how to reason generally. LoRA sidesteps this by leaving the base model untouched and training a small set of additional parameters alongside it — a thin layer of expertise bolted onto a frozen foundation. You get real domain depth without touching the thing that makes the model work at all.

The business logic that follows from this is the actual point, and it’s simpler than it looks:

You stop being hostage to any one model provider — if a better open-weight model ships next month, you port your adapter, not your whole product. You can serve hundreds of differently-customized clients off one base model on one piece of hardware, instead of running a separate giant model per customer. You can ship a fix in an afternoon, because an adapter is a few hundred megabytes, not a training run. And in regulated industries, your proprietary data can train an adapter that never leaves your own infrastructure.

None of this is really a story about model architecture. It’s a story about where the moat moved. For a while the moat was raw capability — whoever had the best model won. Apple and Sierra are betting the moat is now somewhere else entirely: in how tightly you can weave a commodity intelligence into a specific workflow, a specific dataset, a specific customer relationship. The engine is free. The adapter is the business.

Categories
AI

What the Lessor Keeps

Two airlines can fly the same airplane. Not airplanes of the same type — the same airplane, serial number and all, handed back at the end of a lease and reassigned, sometimes within weeks, to a competitor on another continent. AerCap owns more commercial aircraft than any airline on earth, and it leases them to airlines that spend their advertising budgets convincing passengers that flying them is a distinctive experience. The 737 MAX that wears Ryanair’s livery this year might wear Lion Air’s the next, repainted, recertified, its avionics untouched, its airframe indifferent to the change of ownership. The lessor does not care who is flying its asset. It cares that the asset comes back in airworthy condition and that the lease payments clear.

What the airline owns, in the sense that matters, is never the aircraft. It is the route network built up over decades of slot negotiations at constrained airports. It is the maintenance log — every inspection, every part swapped, every anomaly a mechanic in Singapore flagged in 2019 that turned out to predict a fatigue crack nobody else had seen yet. None of that travels with the airplane when the lease ends. It stays behind, compounding, in systems the airline built and the lessor never touches.

Karl Mehta, who has spent a career inside enterprise software watching this kind of asymmetry repeat itself, put a version of it plainly: a model is a brain you rent, and you and your competitor rent the same one. The formulation has the compression of something that has been tested in a few dozen meetings before it found that sentence. It is also, structurally, the airplane story. Anthropic and OpenAI and Google are AerCap. They retain residual value on enormous capital assets — clusters of GPUs depreciating on a schedule, weights trained at a cost that only a handful of balance sheets in the world can absorb — and they lease access to those assets by the token, to anyone who can pay, including, in the same afternoon, two companies trying to put each other out of business. The model does not know whose prompt it is answering. It has no loyalty file. It has, in fact, no memory at all, in the ordinary sense of the word — each call begins exactly where the last one ended for everybody, which is nowhere.

The asymmetry that airlines exploit is the one available here too, and it sits one layer up from the engine. Call it the embedding store, the vector database, the fine-tuning corpus, the retrieval index — the terminology varies by vendor, but the function is constant. It is the accumulated, indexed residue of every customer interaction a company has had, structured so that the rented brain can be handed the relevant fragment of it at the moment of each new call. A bank’s fraud model and a competing bank’s fraud model can call the identical foundation model, route through the identical API, and arrive at entirely different verdicts on the identical transaction, because one of them is retrieving against eleven years of labeled chargebacks specific to its own card portfolio and the other is retrieving against four. The intelligence rented by the hour is, for practical purposes, a commodity, priced down toward marginal cost the way jet fuel is priced — everyone pays close to the same number per unit. The memory is not a commodity. It cannot be, because it is not for sale; it is the institutional record of what has already happened to you, and no amount of capital lets a competitor buy a copy of your chargeback history any more than it lets them buy your maintenance logs.

This produces a particular kind of corporate vertigo, which Mehta’s sentence is really addressing. For three or four years the industry conversation about artificial intelligence has been a conversation about models — which lab’s was larger, which benchmark moved, which release cycle a company should anchor its roadmap to. That conversation rewards being an early and aggressive lessee. But a lessee relationship, however aggressive, does not compound into anything a competitor cannot eventually also lease. The compounding, when it happens, happens in the layer below the API call: in how cleanly a company has structured the record of its own customers, its own failures, its own edge cases, so that the rented brain, plugged in fresh every morning with no memory of yesterday, can be handed exactly the right fragment of yesterday and made to look, for a few hundred milliseconds, like it has been there all along.

A hospital chart has two kinds of entries. There is the vital-signs strip clipped to the bed rail — temperature, pulse, blood pressure, checked every four hours and replaced every four hours, because a reading from yesterday tells the night nurse nothing about the patient in front of her right now. And there is the permanent record in the file downstairs: the allergy that nearly killed him in 2019, the surgery, the medication history going back a decade, written once and never overwritten, because that record is exactly as valuable ten years from now as it is today. Nobody confuses the two charts. Nobody staples last Tuesday’s blood pressure into the permanent file. The hospital figured out, long before anyone digitized it, that memory is not one problem. It is two, and they fail in opposite directions if you run them through the same system.

Most teams building the layer Mehta is describing make exactly that mistake — they staple everything to the same chart. The shorthand for it is dumping everything into a vector database and praying, and it is worth asking why that particular error is so popular. The answer is that it feels like progress: embeddings go in, something resembling memory comes out, and the team moves on to the next sprint without confronting the harder question, which is what kind of memory it just built.

Short-term memory is the vital-signs strip — everything the model needs to finish the task in front of it and nothing it needs after. A customer-service exchange in progress, the order number already mentioned, the fact that this is the second call today, belongs here. So does the scratchpad of a multi-step agent: the search results just pulled, the file just opened, the partial answer being assembled before it commits. The test is not how important the information is but how long it stays true. A customer’s mood this minute is real and gone in twenty minutes; storing it permanently is like stapling yesterday’s temperature reading into the permanent file, undated, until the chart tells you nothing about fever and everything about clutter. Short-term memory should live in the context window itself, or a session-scoped cache, and it should be allowed to die when the session ends. The sin is not forgetting it. The sin is remembering it forever.

Long-term memory is the file downstairs, and it does not come in one shape any more than that file does. The first shape is semantic memory — facts. A customer’s account tier. The chargeback history that decides, in fractions of a second, whether this morning’s transaction clears. Facts belong in a database with a schema, not a vector store, because a fact has a right answer and a vector store gives you an approximate neighbor. Ask a vector index what tier a customer is on and it hands you the five most semantically similar sentences in the corpus — one correct, four merely correct-sounding. Ask a schema the same question and it tells you, because that is what the schema is for.

The more sophisticated shops are already building the seam between the two, rather than picking one and living with its blind spot. A knowledge graph keeps the relationships a schema is good at — this customer, that account, this chargeback, in fixed and queryable connection to one another — while still letting a retrieval layer search across it by meaning rather than by exact key. The approach has a name now, GraphRAG, and the name matters less than what it concedes: that facts and resemblance are different operations, and the honest fix is to run both and let each one answer the kind of question it’s actually suited for, not to force a single index to pretend it can do both jobs at once.

The second shape is episodic memory — what actually happened. The specific conversation last March in which the customer explained, at length, why the previous fix didn’t work. The exact sequence of an agent’s failed attempt at a task, preserved so the next attempt doesn’t repeat it. This is where the vector store finally earns its keep, because an episode isn’t an exact-match lookup, it’s a resemblance — has anything like this come up before — and a vector index, built to find the nearest thing to a fuzzy question, is the right tool for that question and almost no other. The error was never using a vector store. The error is using only a vector store, for facts as well as episodes, on the theory that one hammer with sufficient cosine similarity can stand in for the whole toolbox.

The third shape is the rarest, and the one teams forget to build at all: procedural memory, which is not a fact and not an episode but a skill — the model’s learned sense of how this company writes a refund email, escalates a complaint, formats an invoice. Style is the visible half of it. The other half is harder to see and matters more: the rails the model is forced to run on before it ever gets to choose a word. A refund above some threshold routes to a human, no exceptions, because the workflow says so, not because the model was persuaded to think so on this particular call. An agent that touches a production database does it through a reviewed function with a fixed set of permitted calls, not through whatever query it improvises in the moment. None of that lives in a prompt, and none of it lives in the model’s weights either. It lives in code — the orchestration layer, the permissioning, the state machine the agent is required to pass through — and it is procedural in the oldest sense of the word: not a memory of what to say but a memory of what is and isn’t allowed to happen, enforced whether or not the model that day feels like remembering it. It doesn’t live in a database at all. It lives in fine-tuning, in carefully maintained house-style examples, and in the surrounding scaffolding of guardrails and permitted actions, and it changes slower than the other two, the way a surgeon’s hands carry both technique and caution years after the specific patients are forgotten. A company that has built rich semantic and episodic memory but skipped this layer has a model that knows everything about its customers, writes in exactly the right voice, and is one well-crafted prompt away from doing something the company never agreed to.

The real argument here is not which database serves which layer — that part is plumbing, and plumbing changes every eighteen months. The argument is that memory has to be triaged the way the hospital triages it, with something deciding on purpose what survives the session and what doesn’t, rather than writing every token of every interaction into the same undifferentiated store and trusting retrieval to sort it out later. A vector database with no triage in front of it is not a memory system. It is a landfill with a search function, and it will retrieve the wrong eleven-month-old conversation with the same confidence it retrieves the right one, because nobody wrote the part of the system whose only job is deciding what belongs on which chart.

The lessor’s airplane, repainted, will fly for someone else next year. The route network will not. Neither will the schema that knows a customer’s tier on contact, nor the index that remembers the conversation from last March, nor the fine-tuned hand that knows, without being told twice, how this company writes a refund email. These are the things that do not come back at the end of the lease, because they were never on it.

Categories
AI Consulting

The Judgment Layer

An analyst’s note about the CEO of one of the largest consulting companies making comments at an investor conference includes a line that deserves more attention than it got: “token volume used on a project isn’t a proxy for AI maturity.”

Translation — clients are burning money on frontier models for problems that don’t need frontier models, and they’re not getting the outcomes they expected.

This firm’s CEO offered this as a business opportunity. I read it as a confession.

The old consulting model was simple: client has a technology problem, firm deploys humans to solve it. Billing followed effort. The new problem is different in kind — clients have an AI strategy problem. They know they’re supposed to be using AI. They’ve heard the word “frontier.” They’re spending accordingly. They just don’t know why, and the outcomes are showing it.

So the CEO is right that there’s an opportunity here. The value proposition shifts from implementation to judgment — not deploying AI, but knowing when not to deploy the expensive one. Matching capability to problem. Being trusted enough to tell a client that their $50M frontier model contract is solving a $500K problem.

Here’s the irony that the comment skates past: that advice is structurally difficult for a large consultancy to give.

The business model that built consulting firms was billing for doing. The more you deploy, the more you bill. Helping a client spend less, or choose the cheaper model, or run a narrower project, is genuinely good advice that the incentive structure actively works against. You don’t grow a $70 billion professional services firm by talking clients out of scope.

The judgment layer, if it becomes the real value, requires something closer to a doctor’s relationship with a patient than a contractor’s relationship with a client. Doctors get paid whether they prescribe or not. The value of the visit is the diagnosis — including the diagnosis that says you don’t need the expensive intervention. Consultants, historically, get paid to prescribe, and paid more when the prescription is larger.

There’s a reason we trust doctors with that asymmetry and not contractors. Licensing, malpractice, professional norms built over centuries — all of it exists to align the incentive. Consulting has none of that infrastructure. What it has instead is reputation, which is slower-acting and easier to game.

Whether the large firms can actually make the shift — rather than just reframe the same billable-hours model in the language of AI optimization — is the real question the market is wrestling with. The CEO’s comment is genuinely perceptive about where client value lies. It’s less clear that consulting firms are currently built to capture it honestly.