Categories
AI Business

The Reverse Information Paradox We’ve Always Had

Satya Nadella wrote recently about what he calls the Reverse Information Paradox: enterprises pay for AI intelligence twice. Once in money. Again in the proprietary knowledge they surrender through every prompt, correction, and evaluation. The better they use the model, the more of their own institutional understanding leaks into someone else’s system. The vendor ends up knowing more about the buyer’s business than the buyer knows about what the vendor retained.

Replace “model” with “employee” (or “consultant”) and the paradox is not new at all.

You pay for a person once with salary. You pay again with something harder to price: the context, relationships, and judgment they must absorb to become useful to you. The better they perform, the deeper the immersion, the more of your particular way of doing things moves into their head. Every correction and late-night conversation is another trace of institutional memory changing hands. When they leave, some of that memory leaves with them. Not always through theft. Usually just through the ordinary residue of good work.

The visible cost is salary; the invisible cost is the slow transfer of what makes you distinctive. High performers get more access precisely because they’re high performers, which means the leakage accelerates exactly when you can least afford it. The exhaust is just harder to see with people than with tokens — it moves through conversation and mental models instead of logs.

The analogy has a limit, and the limit matters. Employees bring knowledge in, not just absorb it. They have judgment and relationships a model doesn’t. Models are purely absorptive, and once something is inside them, it’s infinitely reproducible — a person can only be in one place, working for one employer, at a time. We’ve had a few hundred years to build tools for the human version of this problem: contracts, culture, non-competes. The model equivalent is still being invented in real time, which is exactly why Nadella felt the need to name it.

Apple’s recent legal action against former employees who joined OpenAI is this pattern in its sharpest form. Whatever the specifics, the shape is familiar: people who spent years inside one of the most sophisticated organizations in the world, carrying out knowledge that never appeared on any balance sheet and was hard to contain. No one fully anticipates what a mind absorbs simply by being in the room long enough.

That’s the real difference between the silicon case and the human one. You can try to take action to wall off knowledge flowing to a model. You cannot wall off what someone has learned to notice.

Categories
AI

Context Rot

Here is a small, possibly embarrassing confession: I have never, not once, gone looking for the best AI model.

I have a model. It lives in a browser tab — Safari, usually, on whichever device is nearest, occasionally Chrome if I happen to be at the desktop. It does what I need — drafts an email, untangles a sentence, tells me what a Norwegian emigration record from 1856 probably says — and then I close the tab and go on a walk.

Somewhere out there, presumably, a much smarter, much more expensive machine is doing something extraordinary with protein folding or hedge fund arbitrage or the outer edges of mathematics I will never visit. I have made my peace with never meeting it.

This did not used to feel like a confession. For a while there — a year, eighteen months — it felt like the central drama of the whole industry: which model was “best,” who had it, who had lost it, whether some lab’s quarterly earnings call would reveal that the frontier had quietly moved sixty miles down the road while everyone was looking the other way. Benchmarks were released like box scores. People argued about them the way people argue about batting averages, with the same weird intensity, the same conviction that a two-point difference in some abstract reasoning test settled something important about the future.

And then, at some point I can’t quite date — it crept up, the way these things do — I noticed I had stopped caring.

Not because the frontier stopped moving. It didn’t. It’s still moving, arguably faster than ever, in ways that occasionally show up in the news with all the drama of a soap opera (a delayed launch, a researcher poached, a stock down five percent in an afternoon, always something).

I stopped caring because none of it touched me. My model — whatever it was, this week — had long since crossed some invisible threshold past which more didn’t register as more. It was already better than I needed. It has been better than I needed for a while now. I suspect I am not unusual in this. I suspect most people, doing most things, most days, are operating comfortably inside a capability surplus so large they’ve stopped noticing it’s there, the way you stop noticing a room is warm.

If the top of the model isn’t for people like me — and it increasingly isn’t — then who, or what, is it actually for? I went looking for one piece of the answer and found, instead, a metaphor.

It’s called “context rot.” I have to admit, before I go further, that I’m not sure I’ve ever felt it myself — which, on reflection, is its own small piece of evidence. My sessions close in minutes, not hours. I ask, it answers, I leave. Whatever happens to a model over the fourth or fifth hour of sustained, dependent work is a country I simply don’t visit.

But other people do, increasingly — entire teams do, for entire projects — and what they’re finding out there is worth understanding, even secondhand. It describes something that happens to AI models when they’re asked to work for a long time on something complicated — not five minutes, but five hours; not one question, but a hundred small decisions stacked on top of each other, each one depending on the last.

You’d think the limiting factor would be room. Models have a “context window” — a stated capacity, like a gas tank, measured in tokens, and for a while the marketing numbers on these were the whole story: two million tokens! A library! And you’d think, as with a gas tank, that the thing runs fine until it’s empty and then it stops.

That is not, it turns out, what happens. What happens is closer to what happens to your desk.

You know the desk. Everyone has the desk. It starts the morning clean — an aspirational, almost insulting cleanliness — and by four in the afternoon it is a geological record of the day: three coffee cups, a stack of things you meant to file, a Post-it with a phone number you no longer need, the good pen buried under a printout of something you already dealt with an hour ago. The desk is not full. There is, technically, room. You could clear a space if you tried. But you don’t try, because functionally, cognitively, the desk has stopped being usable long before it ran out of surface area. You start looking for the stapler and forget what you were stapling. This — and I did not make this term up, I want to be clear, though I wish I had — is context rot. The window hasn’t run out. The signal has just drowned in its own debris.

Researchers watching this happen to long-running AI agents have found something almost cruelly elegant about how it fails: it doesn’t fail gradually, the way you’d expect a desk to get gradually messier. Errors compound. A task that takes twice as long doesn’t get twice as likely to go wrong — the failure rate roughly quadruples. Two mistakes early in a long chain of dependent steps don’t add up to a slightly worse outcome. They multiply into something close to total collapse, four hours in, for reasons that trace back to a single bad assumption made in the first twenty minutes and never revisited.

Here is where the frontier comes back in — not as the whole answer, but as a piece of one.

It is not that frontier models are smarter in the way a benchmark measures smart — better at a single hard math problem, a cleverer turn of reasoning. Plenty of models can do that now; the “good enough” tier has crept remarkably high.

It’s that frontier models are apparently, marginally, meaningfully better at not rotting. At keeping the desk usable at hour six. At knowing which of the forty things on the desk actually still matters and which is a coffee cup that should have been thrown out an hour ago. This is a genuinely different kind of intelligence than the one benchmarks were built to measure, and it is almost invisible from the outside — you don’t see it in a single exchange, you see it only in the difference between a project that holds together over three days and one that quietly, subtly, stops making sense somewhere around Tuesday afternoon and nobody notices until Thursday.

If that’s true — if the frontier’s real edge is durability rather than raw cleverness — you’d expect to see it show up in how the labs actually deploy their own models: saving the sharpest tools for the tasks that need to survive the longest.

I went looking for a real-world example and found one closer to home than I expected: Anthropic’s own Slack tool, the one where you tag the AI into a channel the way you’d tag a coworker, and it works alongside a whole team over days, learning the channel as it goes. It runs on a serious, capable, thoroughly frontier model — but not, it turns out, on the company’s very best one. That one is held back, reserved for a smaller and stranger set of problems nobody has solved before at all. I sat with that for a while. The tool built to survive a whole team’s whole week, in public, under the most sustained pressure any of their products face, wasn’t handed the sharpest blade in the drawer. It was handed the second-sharpest — which was apparently, entirely, enough. Which tells you something about where the two kinds of intelligence actually diverge: the merely-very-good model handles the desk staying clean for a week, in public, in front of a whole team, where one bad assumption made Monday and never revisited would be visible to everyone by Thursday. The truly new capability is being held in reserve for something else altogether.

I don’t have a tidy place to land this, and I’m suspicious of anyone who does. But here’s the closest I can get.

Imagine a three-Michelin-star chef — the kind of person who has spent thirty years learning to coax something transcendent out of a single scallop, who can tell you, by smell, that a stock has forty more minutes in it — standing at your stove on a Tuesday night making you a grilled cheese sandwich. It will, I promise you, be a very good grilled cheese sandwich. The bread will be evenly golden. The cheese will have reached some ideal, fully-considered state of melt. But almost none of what makes that chef extraordinary is actually being used to make it — none of the thirty years spent learning to hold forty things in mind at once without losing track of any of them, the exact skill, it occurs to me, that keeps a long, complicated project from quietly falling apart on day three. The technique is idling. The thirty years are in the room, present, available, and almost entirely beside the point, because a grilled cheese sandwich was never the place where thirty years shows up. It shows up somewhere else — in a dish you will never order, on a night you weren’t there.

What you got instead, on your ordinary Tuesday, was simply more than enough.

Categories
AI Learning Photography

Autopilot

“Superb photographs are not just taken with cameras. They come from within you, your eyes, your mind, your heart, not ice cold equipment.” Fan Ho

There’s a half-second on the street, somewhere between seeing a frame and shooting it, that used to take me whole minutes. Early on, with a camera in my hands on the streets of San Francisco or on the subway platforms in New York, I’d see something — light falling a certain way, a gesture about to resolve into a gesture — and I’d think my way through it. Assess the composition or the angle. Worry about the background. By the time I’d worked it out, the moment might be gone, replaced by some lesser version of itself.

That doesn’t happen to me anymore, and I couldn’t tell you when it stopped. Somewhere along the way the thinking disappeared and the shooting stayed. I see the frame and the shutter goes, and only afterward, looking at the file, do I understand what I saw. I didn’t explicitly decide to skip the thinking. It just stopped showing up, the way a habit eventually stops asking your permission. Or how driving a car becomes second nature.

I think about this because of a problem the AI labs have been calling continual learning. The AI models we use are like brilliant interns. They can solve a hard problem at nine in the morning and a harder one by five, and they’ll astonish you doing it. But every session starts over from zero. Whatever they got right on Tuesday evaporates by Wednesday, the way a dream is gone by the time you’ve found your slippers.

The industry’s first answer was to give them a longer memory — let the window hold the whole case file in front of them, all the time. This works for a while, the same way it would work for me on the street if I stopped and re-derived the exposure math for every frame. But that isn’t how I shoot anymore. I don’t have the math open. I have what’s left after thousands of frames did the math for me and then got out of the way.

Based on some exploration I did this morning using AI I found three different AI research efforts that are now chasing that gap, from different angles, none of them all the way there.

A team out of Stanford and NVIDIA built something called TTT-E2E, which lets a model keep adjusting its own internal weights while it reads — not just holding the page in front of it, but being changed by the page, a little, as it goes. It runs thirty-five times faster than the brute-force method of remembering everything, because it isn’t remembering everything.

Google’s research arm published something called Nested Learning around the same time, built on the idea that a mind isn’t one system learning at one speed, but several systems nested inside each other — some updating by the minute, some by the year.

And a scrappier strand of work called self-distillation has models teaching cheaper versions of themselves, not by handing over a transcript, but by training the cheaper model to arrive on its own at whatever the well-informed version would have concluded.

None of this is what happens when I make a photo. Not yet. But it’s aimed at the same gap I live in every time I shoot before I understand what I’m shooting. The gap between having the math and having the eye.

I once asked Doug, a good friend who’s spent as many days on the street as I have, how he knew when to press the shutter. He didn’t have an answer, not really — just a shrug, and something about the moment feeling complete before he could explain why. That shrug took him years to earn. He didn’t keep the years. He kept the shrug.

And then a few years ago Doug did something I still don’t fully understand. He abandoned digital and went back to film. Not for any project, not for the look of it — he could get that in post if he wanted it. He went back to the actual mechanics: loading a roll, metering by hand, often using a tripod, etc. I needled him about it some, the way you’d needle a cigarette smoker who’d taken up a pipe instead, as if the inconvenience were the point. He told me he wanted to slow down, and that film was the only thing that reliably made him do it. Twelve frames and then you stop and reload and you can’t fix it later. The very friction he’d spent decades shooting his way out of, he went looking for again, on purpose.

I don’t know what to do with that, except to notice that he’s the same man who can give me the shrug and also the man who walked back toward the thing the shrug had replaced. Maybe that’s the part the labs haven’t gotten to yet, underneath all the vocabulary of weight updates and meta-learned initializations. Compression is the whole point, until the day it isn’t.

Note: This line of thinking started with a recent essay by Dwarkesh Patel on what he calls continual learning. It’s become a real focus of his thinking about how we get to a better future with AI.

See: https://www.dwarkesh.com/p/the-next-paradigm

Categories
AI

What the Lessor Keeps

Two airlines can fly the same airplane. Not airplanes of the same type — the same airplane, serial number and all, handed back at the end of a lease and reassigned, sometimes within weeks, to a competitor on another continent. AerCap owns more commercial aircraft than any airline on earth, and it leases them to airlines that spend their advertising budgets convincing passengers that flying them is a distinctive experience. The 737 MAX that wears Ryanair’s livery this year might wear Lion Air’s the next, repainted, recertified, its avionics untouched, its airframe indifferent to the change of ownership. The lessor does not care who is flying its asset. It cares that the asset comes back in airworthy condition and that the lease payments clear.

What the airline owns, in the sense that matters, is never the aircraft. It is the route network built up over decades of slot negotiations at constrained airports. It is the maintenance log — every inspection, every part swapped, every anomaly a mechanic in Singapore flagged in 2019 that turned out to predict a fatigue crack nobody else had seen yet. None of that travels with the airplane when the lease ends. It stays behind, compounding, in systems the airline built and the lessor never touches.

Karl Mehta, who has spent a career inside enterprise software watching this kind of asymmetry repeat itself, put a version of it plainly: a model is a brain you rent, and you and your competitor rent the same one. The formulation has the compression of something that has been tested in a few dozen meetings before it found that sentence. It is also, structurally, the airplane story. Anthropic and OpenAI and Google are AerCap. They retain residual value on enormous capital assets — clusters of GPUs depreciating on a schedule, weights trained at a cost that only a handful of balance sheets in the world can absorb — and they lease access to those assets by the token, to anyone who can pay, including, in the same afternoon, two companies trying to put each other out of business. The model does not know whose prompt it is answering. It has no loyalty file. It has, in fact, no memory at all, in the ordinary sense of the word — each call begins exactly where the last one ended for everybody, which is nowhere.

The asymmetry that airlines exploit is the one available here too, and it sits one layer up from the engine. Call it the embedding store, the vector database, the fine-tuning corpus, the retrieval index — the terminology varies by vendor, but the function is constant. It is the accumulated, indexed residue of every customer interaction a company has had, structured so that the rented brain can be handed the relevant fragment of it at the moment of each new call. A bank’s fraud model and a competing bank’s fraud model can call the identical foundation model, route through the identical API, and arrive at entirely different verdicts on the identical transaction, because one of them is retrieving against eleven years of labeled chargebacks specific to its own card portfolio and the other is retrieving against four. The intelligence rented by the hour is, for practical purposes, a commodity, priced down toward marginal cost the way jet fuel is priced — everyone pays close to the same number per unit. The memory is not a commodity. It cannot be, because it is not for sale; it is the institutional record of what has already happened to you, and no amount of capital lets a competitor buy a copy of your chargeback history any more than it lets them buy your maintenance logs.

This produces a particular kind of corporate vertigo, which Mehta’s sentence is really addressing. For three or four years the industry conversation about artificial intelligence has been a conversation about models — which lab’s was larger, which benchmark moved, which release cycle a company should anchor its roadmap to. That conversation rewards being an early and aggressive lessee. But a lessee relationship, however aggressive, does not compound into anything a competitor cannot eventually also lease. The compounding, when it happens, happens in the layer below the API call: in how cleanly a company has structured the record of its own customers, its own failures, its own edge cases, so that the rented brain, plugged in fresh every morning with no memory of yesterday, can be handed exactly the right fragment of yesterday and made to look, for a few hundred milliseconds, like it has been there all along.

A hospital chart has two kinds of entries. There is the vital-signs strip clipped to the bed rail — temperature, pulse, blood pressure, checked every four hours and replaced every four hours, because a reading from yesterday tells the night nurse nothing about the patient in front of her right now. And there is the permanent record in the file downstairs: the allergy that nearly killed him in 2019, the surgery, the medication history going back a decade, written once and never overwritten, because that record is exactly as valuable ten years from now as it is today. Nobody confuses the two charts. Nobody staples last Tuesday’s blood pressure into the permanent file. The hospital figured out, long before anyone digitized it, that memory is not one problem. It is two, and they fail in opposite directions if you run them through the same system.

Most teams building the layer Mehta is describing make exactly that mistake — they staple everything to the same chart. The shorthand for it is dumping everything into a vector database and praying, and it is worth asking why that particular error is so popular. The answer is that it feels like progress: embeddings go in, something resembling memory comes out, and the team moves on to the next sprint without confronting the harder question, which is what kind of memory it just built.

Short-term memory is the vital-signs strip — everything the model needs to finish the task in front of it and nothing it needs after. A customer-service exchange in progress, the order number already mentioned, the fact that this is the second call today, belongs here. So does the scratchpad of a multi-step agent: the search results just pulled, the file just opened, the partial answer being assembled before it commits. The test is not how important the information is but how long it stays true. A customer’s mood this minute is real and gone in twenty minutes; storing it permanently is like stapling yesterday’s temperature reading into the permanent file, undated, until the chart tells you nothing about fever and everything about clutter. Short-term memory should live in the context window itself, or a session-scoped cache, and it should be allowed to die when the session ends. The sin is not forgetting it. The sin is remembering it forever.

Long-term memory is the file downstairs, and it does not come in one shape any more than that file does. The first shape is semantic memory — facts. A customer’s account tier. The chargeback history that decides, in fractions of a second, whether this morning’s transaction clears. Facts belong in a database with a schema, not a vector store, because a fact has a right answer and a vector store gives you an approximate neighbor. Ask a vector index what tier a customer is on and it hands you the five most semantically similar sentences in the corpus — one correct, four merely correct-sounding. Ask a schema the same question and it tells you, because that is what the schema is for.

The more sophisticated shops are already building the seam between the two, rather than picking one and living with its blind spot. A knowledge graph keeps the relationships a schema is good at — this customer, that account, this chargeback, in fixed and queryable connection to one another — while still letting a retrieval layer search across it by meaning rather than by exact key. The approach has a name now, GraphRAG, and the name matters less than what it concedes: that facts and resemblance are different operations, and the honest fix is to run both and let each one answer the kind of question it’s actually suited for, not to force a single index to pretend it can do both jobs at once.

The second shape is episodic memory — what actually happened. The specific conversation last March in which the customer explained, at length, why the previous fix didn’t work. The exact sequence of an agent’s failed attempt at a task, preserved so the next attempt doesn’t repeat it. This is where the vector store finally earns its keep, because an episode isn’t an exact-match lookup, it’s a resemblance — has anything like this come up before — and a vector index, built to find the nearest thing to a fuzzy question, is the right tool for that question and almost no other. The error was never using a vector store. The error is using only a vector store, for facts as well as episodes, on the theory that one hammer with sufficient cosine similarity can stand in for the whole toolbox.

The third shape is the rarest, and the one teams forget to build at all: procedural memory, which is not a fact and not an episode but a skill — the model’s learned sense of how this company writes a refund email, escalates a complaint, formats an invoice. Style is the visible half of it. The other half is harder to see and matters more: the rails the model is forced to run on before it ever gets to choose a word. A refund above some threshold routes to a human, no exceptions, because the workflow says so, not because the model was persuaded to think so on this particular call. An agent that touches a production database does it through a reviewed function with a fixed set of permitted calls, not through whatever query it improvises in the moment. None of that lives in a prompt, and none of it lives in the model’s weights either. It lives in code — the orchestration layer, the permissioning, the state machine the agent is required to pass through — and it is procedural in the oldest sense of the word: not a memory of what to say but a memory of what is and isn’t allowed to happen, enforced whether or not the model that day feels like remembering it. It doesn’t live in a database at all. It lives in fine-tuning, in carefully maintained house-style examples, and in the surrounding scaffolding of guardrails and permitted actions, and it changes slower than the other two, the way a surgeon’s hands carry both technique and caution years after the specific patients are forgotten. A company that has built rich semantic and episodic memory but skipped this layer has a model that knows everything about its customers, writes in exactly the right voice, and is one well-crafted prompt away from doing something the company never agreed to.

The real argument here is not which database serves which layer — that part is plumbing, and plumbing changes every eighteen months. The argument is that memory has to be triaged the way the hospital triages it, with something deciding on purpose what survives the session and what doesn’t, rather than writing every token of every interaction into the same undifferentiated store and trusting retrieval to sort it out later. A vector database with no triage in front of it is not a memory system. It is a landfill with a search function, and it will retrieve the wrong eleven-month-old conversation with the same confidence it retrieves the right one, because nobody wrote the part of the system whose only job is deciding what belongs on which chart.

The lessor’s airplane, repainted, will fly for someone else next year. The route network will not. Neither will the schema that knows a customer’s tier on contact, nor the index that remembers the conversation from last March, nor the fine-tuned hand that knows, without being told twice, how this company writes a refund email. These are the things that do not come back at the end of the lease, because they were never on it.