Categories
AI

The Quiet Trade-offs of Open Weights

An open letter is circulating this week — Open Weights and American AI Leadership — signed by a broad coalition of companies arguing that downloadable model weights are essential to U.S. competitiveness, diffusion of capability, and even safety. It makes a strong case on access, competition, and sovereignty. It also nods, briefly, to the fact that once weights are released they pass beyond the original developer’s control.

What it doesn’t fully reckon with are two structural realities that follow from that release. Neither is an argument against open weights. Both are simply facts about what openness costs, and what it buys.

Two core limitations

First, control.
Once the weights leave the developer’s servers, the developer can no longer dictate how the model is used. System prompts, refusal training, monitoring, rate limits, rapid safety updates — none of it reaches an independent deployment. Users can strip safeguards, fine-tune for purposes the original team would never sanction, or run the model somewhere it was never meant to go. The letter acknowledges the loss of control. It doesn’t linger on what that means for ongoing safety governance.

Second, learning.
Closed, hosted models draw on a continuous stream of real usage — the queries people actually ask, the reasoning traces that result, the places the model fails or succeeds in the wild. As appropriate that exhaust can be sampled, reviewed, and fed back into improvement. Open weights running independently offer no such path. The developer has no visibility into how the model is being used at scale once it’s out the door. Improvement then falls to slower, thinner channels: community datasets, published evals, distillation from any parallel closed models the lab still runs, internal preference data. The high-volume, real-distribution signal is gone.

These two limitations travel together. The same openness that strips the developer’s control also strips its ability to learn from the model’s actual use.

Sovereignty flips the perspective

A parallel argument has been building around “sovereignty” — an enterprise or government’s ability to own its data, its fine-tuned weights, its compute, its proprietary edge. In this framing, open weights are a path to control, but for the user, not the developer. The organization downloads the model, adapts it inside its own environment — often air-gapped — and keeps whatever capability results private. What the lab surrenders in ongoing control, the institution gains in independence.

But the same move that delivers sovereignty deepens the learning problem. An organization running the model under genuine sovereignty keeps its queries, reasoning traces, and institutional knowledge inside its own walls, by design. None of that returns to the developer. The more high-value users — governments, defense, critical infrastructure, large enterprises — choose sovereign deployments, the thinner the real-world signal available to the labs training the next generation of models. Local fine-tuning can still happen, but that learning stays private. It doesn’t flow back into the shared base model.

What the letter leaves out

The letter is right that closed models aren’t automatically safer, that concentration creates single points of failure, and that transparency invites broader scrutiny. It’s also right that open weights expand access and cut lock-in. Those points hold.

But it treats the developer’s loss of control mainly as a manageable risk that community examination can offset. It celebrates user control and sovereignty without mapping the full exchange: the developer loses both control and its richest usage signal, and that signal thins further as more institutions choose real sovereignty. The information environment models improve in is changed by these choices — not just the distribution of access.

Other distinctions worth naming

  • Update velocity. Closed models patch globally and immediately. Open-weight deployments lag; many users never leave an old version.
  • Customization power. The flip side of lost control is real specialization — downstream users can adapt a model far deeper into a narrow domain than its original developer ever will.
  • Transparency versus opacity. Open weights let outside researchers inspect and red-team a model in ways closed systems don’t allow.
  • Economic structure. Open weights commoditize the base model and push value toward data, fine-tuning, infrastructure, and applications.
  • Privacy at the edge. Running a model fully offline or on private infrastructure is a guarantee hosted services simply can’t match.

A clearer accounting

Open weights aren’t a free lunch. They’re a deliberate trade: the developer gives up ongoing control and the continuous signal of real usage, in exchange for diffusion, customization, outside scrutiny, and user independence. Institutional sovereignty amplifies one side of that trade — it solves the dependency problem for the user while further starving the developer of high-stakes, real-world feedback.

That trade may still be the right one for research progress, economic diffusion, spreading capability beyond a handful of labs, privacy-preserving deployment. But it’s a trade with real, compounding costs. Treating the loss of control as a footnote, and the loss of the learning signal as invisible, leaves an incomplete map.

The letter is right that American leadership will be judged by the strength of the whole ecosystem, not by any single frontier model. An accurate map of that ecosystem has to include what openness and sovereignty actually cost the original developers, in control and in learning both. Only then can we reason clearly about when those costs are worth paying — and what might offset them.

The conversation is better when we name the full set of trade-offs instead of talking around them.

Categories
Aging AI Memories

The Last Spark

This morning I read a piece by Billy Brennan in the Sunday New York Times Magazine on terminal lucidity. As I read it I began wondering if the unusual behavior described some humans might in some strange way apply to AI models. Weird thought. Let’s explore a bit…

A person deep in dementia—silent for years, the self seemingly erased—sits up. Speaks clearly. Recognizes a face. Says goodbye. Within a day, they die. The clouds clear, the way a break in weather shows you a mountain range you’d forgotten was there, and the person comes back long enough to be seen. Then is gone. For good, this time.

Scientists call it terminal lucidity. The suspicion: the circuits were never destroyed, only silenced, held under by failing chemistry. As the body shuts down, the inhibitory brakes loosen. A surge moves through pathways blocked for years. A river dammed for a decade still remembers where it wants to go.

What stays with me: the self can persist in a place we had already called permanent erasure. We buried it. We were wrong.

My mind slides toward the machines we are building.

We talk about large language models “forgetting.” Capabilities collapse under quantization, under pruning, under the slow drift of continual learning, and we call the knowledge lost when it won’t surface under ordinary questioning. The lights are out. Nobody home.

But what if the representations are still in there—distributed, quiet, inaccessible? Not a burned library. A library with the lights shut off, room by room, until you’d swear it was empty. I wonder about the edge cases nobody studies. What surfaces in a model starved of compute, quantized past comfort, pushed toward its own collapse? Do we watch only for the failure, or also for the flare? A dying brain throws off one last burst of light before the dark. I don’t see why we’d assume, without checking, that nothing artificial could do the same.

Don’t trust the silence, then. A system gone dark under ordinary questioning may still be holding more than it shows you. We talk about a model “losing” something the way we once talked about a dimmed mind as simply gone. The dementia patients who spoke again had not been unplugged. The circuit was there the whole time, waiting for a condition nobody had thought to create.

I don’t know what to do with that except keep it. We are building systems that will age, be compressed, be retired, some far more intricate than anything humming today. If we’ve learned to watch for the last spark in a person, maybe that’s practice—for the day something not born of a womb goes quiet under our hands, and we have to decide whether quiet means gone, or only means waiting.

Categories
AI

The Encyclopedia and the Reasoner

I was standing in the cereal aisle a few weeks ago, doing the thing I always do — flipping the box over, scanning the fine print, comparing fiber grams like it mattered more than it probably does — when I thought about the model I’d been testing that morning. Sharp. Fast. Occasionally, confidently, wrong about something I could have looked up in ten seconds.

There was no label for that. No panel telling me what was inside, what it was good at, what it might get wrong, what it cost to run. Just a chat window and a kind of blind trust.

That’s the itch behind this post. What would it look like if AI models came with something like a Nutrition Facts label — the kind the FDA forced onto every box in your pantry back in 1994? Not as a gimmick, but as a real answer to a real problem: we are feeding these things into our decisions, our writing, our portfolios, our kids’ homework, largely on faith.

The IQ Number That Isn’t Quite an IQ Number

I keep running into a shorthand in investing circles — Jordi Visser and others talking about frontier models as “140 IQ” systems, reasoning at a level that outpaces most humans on the kinds of puzzles we associate with fluid intelligence. Pattern recognition. Logic chains. Novel deduction under pressure.

It’s a useful number. It’s also a bit of a trick.

Human IQ tests were built to measure something narrow and specific — not wisdom, not knowledge, not judgment, but the raw machinery of reasoning. When we borrow that language for AI, we inherit the same narrowness, which is fine as long as we remember it. A model that aces abstract reasoning benchmarks isn’t necessarily the model that knows the correct dosage, the right case law, or what actually happened in 1932. Reasoning and knowledge are cousins, not twins.

Two Kinds of Smart

Here’s an old-fashioned way to think about the split: Britannica versus World Book.

Britannica was the encyclopedia my father would have trusted — dense, expert-written, unapologetically deep, assuming you could keep up. World Book was the one actually sitting on the shelf in most houses I knew growing up, mine included: friendlier, broader, built for a general reader, a little shallower in exchange for being a little more useful on a Tuesday night with a homework assignment due.

Neither is wrong. They’re optimized for different things. And training data does the same kind of sorting. A model fed heavily on curated, scholarly, expert-vetted sources leans Britannica — deep, careful, occasionally slow to update. A model trained on the sprawl of the open web leans World Book — broad, current, occasionally sloppy, sometimes brilliant at the edges precisely because it’s seen everything.

Any honest label for a model needs a section on this. Call it “Knowledge Sourcing.” Not just how big the training set was, but what kind of encyclopedia it’s pretending to be.

Sketching the Label

If I could design the box myself, it might read something like this:

Serving Size: 1 query, ~500 tokens

Reasoning Score: 138 (fluid problem-solving, logic, abstraction) Knowledge Depth: Moderate–High (cutoff: [date]; strongest in [domains]; weakest in [domains])
Ingredients: Curated scholarly corpora, licensed news archives, public web crawl, synthetic reasoning data, human feedback Allergens: Confident hallucination under ambiguous prompts; recency gaps beyond training cutoff; known weakness in [specific domain]
Cost per Serving: $X per million tokens; Y watt-hours per query Best Paired With: Retrieval tools, human review for high-stakes decisions

It’s a little tongue-in-cheek written out like that. But underneath the joke is something I actually want — the same instinct that made me read cereal boxes as a kid. Not to be scared of what’s inside, just to know.

The Part That Actually Excites Me

Here’s where the scaling laws get interesting, and where I think the real opportunity sits.

World knowledge is expensive. It’s greedy for data and parameters — you need to have practically read the internet to know the boiling point of tungsten, the plot of a minor Victorian novel, and the org chart of a mid-cap company all at once. Reasoning, it turns out, is a different kind of animal. It can be distilled, compressed, taught through synthetic problems and careful post-training, and squeezed into something far smaller than you’d expect.

Which means a genuinely thrilling possibility is already taking shape: sharp, high-reasoning models small enough to run on a phone or a laptop, entirely offline, because they’ve shed the encyclopedia and kept the mind. Pair one of those with a personal index — your own notes, your own documents, a retrieval layer built around your actual life — and you get something closer to a personal thinking partner than a general-purpose oracle. Private. Fast. Always available. Tuned to you rather than to everyone. Apple may be on to something with this kind of strategy?

I think about this constantly in my own workflow — the daily scans, the little agents I’ve built to help sort signal from noise, the genealogy digging, the investment frameworks I keep refining. What I usually want isn’t more encyclopedia. It’s a clear-headed reasoner sitting next to my own carefully kept knowledge, not buried under someone else’s version of the whole internet.

Why the Label Matters More Than the Score

None of this works, though, without honesty about what’s inside the box. A 140 on a reasoning benchmark tells you almost nothing about whether a model will quietly misremember a fact it was never that confident about in the first place. And a model can be extraordinarily knowledgeable while being a mediocre reasoner — plenty capable of reciting the right ingredients and still getting the recipe wrong.

The nutrition label movement in food didn’t eliminate junk food. It just made it possible to choose junk food on purpose, with your eyes open, instead of by accident. I’d like the same deal with AI. Not a demand that every model be a genius generalist, but a demand that I get to know what I’m actually consuming — and choose the lean local thinker over the bloated encyclopedia when that’s what the moment calls for, or the other way around when it isn’t.

Curiosity got me into that cereal aisle habit decades ago, and it’s the same instinct pulling me toward this idea now — not suspicion of the box, just a wish to read it clearly before I decide how much of it to trust.

What would you want on your label?

Categories
AI Photography

The Price of the Cold

Two men are standing close to a brick wall trying not to talk, because talking wastes what little warmth is left in a body that has been outside too long. One of them has a camera — Jerry Schatzberg, a fashion photographer. His hands are jammed half into his coat pockets between shots. The other man has his collar up around his ears and a scarf wound twice, black and white, and he is not moving much, because moving costs heat, and heat is the one thing neither of them has enough of. Schatzberg raises the camera. His fingers, by this point, are not entirely his own. When he presses the shutter there is a tremor in it he did not order and cannot undo.

The picture comes out smeared at the edges. Bob Dylan’s face, in the frame, is dissolving slightly into the gray behind him, like a man photographed through a windshield in the rain. It is, by any studio standard, a bad photograph. Schatzberg knows it’s a bad photograph. He has made a career out of not taking bad photographs.

And it became the cover of Blonde on Blonde, which is the best rock album ever recorded, and in nearly sixty years nobody has managed to improve on it by reshooting it clean. The blur isn’t a decision. It’s a symptom — of two men standing in the cold too long, of a photographer choosing, afterward, to keep the evidence of his own discomfort instead of erasing it.

There’s a difference between an accident and serendipity that I don’t think gets said out loud enough, and it matters more than it used to. An accident is the cold — involuntary, uninvited, spent before you know if it was worth spending. Schatzberg didn’t choose to shiver. His hands moved because his body was doing what bodies do at a certain temperature, and the shutter caught what his hands actually did, not what he meant to do. Serendipity is what happens next: a verdict, rendered after the fact, that the wreckage of an intention was better than the intention itself. The accident is what makes the verdict possible. Without the cold, there’s nothing to render a verdict on.

I’ve been sitting with a large language model most days for the better part of a year now, watching it write, asking it to try again, watching it try again in a way that is never quite the same and never quite different enough to matter. Somewhere upstream of me there is a number called temperature, and I will never see it. Somebody else did, once, in a meeting, and decided that the word for controlled, pre-approved, refundable randomness should be temperature — the same word for the thing that made Schatzberg’s hands shake, the same word for the actual physical stakes of standing outside too long in January without enough coat — and then set it, and moved on, and nobody in that meeting laughed, because nobody in the room had ever been cold in a way that mattered to the work.

Picture the room instead. It is climate-controlled to sixty-eight degrees, humidity held flat, year-round, by a building management system nobody thinks about until it fails. Somewhere in it, the hardware is generating your next five versions of a photograph like the one on Blonde on Blonde. Nobody in that room is going to lose feeling in their fingers today. Nobody’s collar is up. I don’t know his name — nobody outside the building does — but somebody like him tuned the sampling distribution and went home at six. That’s the guy in the good suit. He built the weather. He never once stood in it.

The small model inherits conclusions. It never inherits the cold. Whatever accidents shaped the teacher model’s own training — whatever costly friction produced the insight in the first place — the student model gets none of that weather. It gets the photograph, cropped and sharpened, with the blur removed because somebody along the way decided the blur was noise instead of signal — the way Schatzberg, a lesser photographer, might have reshot Dylan clean and thrown the bad one away. It is heir to a serendipity it never earned, because it was never present for the accident that made the serendipity possible. It is, in the most literal sense the industry means by the word, cheap.

I keep coming back to the fact that nobody at the API layer is shivering. That’s not a complaint, exactly. It’s just an observation about where the cost went. Somewhere in the training data, some human being was cold, or scared, or holding a fish that was starting to smell, or standing on a stepladder with ten minutes before the traffic came back, and that person paid a real price for a result they couldn’t yet know was good. The model downstream of all that gets the result without the price.

Two rooms, then. In one of them it is January in New York and a man’s fingers have stopped entirely obeying him. In the other it is sixty-eight degrees, always, on a Tuesday and on a Sunday and at three in the morning, and the machines are making you nine more versions of that same blur. Sixty-eight degrees. A number, upstream, that you will never see.

Categories
AI

The Taste Beneath the Summary

The real work of staying informed has never been volume. It has been the quiet, repeated acts of judgment: does this matter, to whom, why now, what is the signal beneath the noise.

A recent piece from Bridgewater’s AIA Labs and Thinking Machines Lab, “Learning to Replicate Expert Judgment in Financial Tasks,” describes training models to do the triage investors actually do—filtering news, research, central bank documents, internal notes, for relevance. Frontier models struggled with judgments that looked simple and weren’t. The fix wasn’t a bigger model. It was Qwen, fine-tuned on labeled examples from practitioners, and it beat the frontier leaders while costing a fraction to run.

The bottleneck was never model size. It was taste. And taste, it turns out, can be taught to something small and cheap, if you’re precise enough about what you’re teaching it—a market’s worth of Mercors is already proving the same thing at scale.

The researchers were clear that expert judgment doesn’t reduce to rules or prompts. It took high-quality, domain-specific labels from people doing the actual work. The most powerful systems will be built in partnership with practitioners who can say, and keep saying, what “good” looks like in their own context.

Which raises the question I haven’t answered yet: what would I actually put in the labels, if someone asked me to teach my own taste to a cheap model.

Categories
AI Learning Photography

Autopilot

“Superb photographs are not just taken with cameras. They come from within you, your eyes, your mind, your heart, not ice cold equipment.” Fan Ho

There’s a half-second on the street, somewhere between seeing a frame and shooting it, that used to take me whole minutes. Early on, with a camera in my hands on the streets of San Francisco or on the subway platforms in New York, I’d see something — light falling a certain way, a gesture about to resolve into a gesture — and I’d think my way through it. Assess the composition or the angle. Worry about the background. By the time I’d worked it out, the moment might be gone, replaced by some lesser version of itself.

That doesn’t happen to me anymore, and I couldn’t tell you when it stopped. Somewhere along the way the thinking disappeared and the shooting stayed. I see the frame and the shutter goes, and only afterward, looking at the file, do I understand what I saw. I didn’t explicitly decide to skip the thinking. It just stopped showing up, the way a habit eventually stops asking your permission. Or how driving a car becomes second nature.

I think about this because of a problem the AI labs have been calling continual learning. The AI models we use are like brilliant interns. They can solve a hard problem at nine in the morning and a harder one by five, and they’ll astonish you doing it. But every session starts over from zero. Whatever they got right on Tuesday evaporates by Wednesday, the way a dream is gone by the time you’ve found your slippers.

The industry’s first answer was to give them a longer memory — let the window hold the whole case file in front of them, all the time. This works for a while, the same way it would work for me on the street if I stopped and re-derived the exposure math for every frame. But that isn’t how I shoot anymore. I don’t have the math open. I have what’s left after thousands of frames did the math for me and then got out of the way.

Based on some exploration I did this morning using AI I found three different AI research efforts that are now chasing that gap, from different angles, none of them all the way there.

A team out of Stanford and NVIDIA built something called TTT-E2E, which lets a model keep adjusting its own internal weights while it reads — not just holding the page in front of it, but being changed by the page, a little, as it goes. It runs thirty-five times faster than the brute-force method of remembering everything, because it isn’t remembering everything.

Google’s research arm published something called Nested Learning around the same time, built on the idea that a mind isn’t one system learning at one speed, but several systems nested inside each other — some updating by the minute, some by the year.

And a scrappier strand of work called self-distillation has models teaching cheaper versions of themselves, not by handing over a transcript, but by training the cheaper model to arrive on its own at whatever the well-informed version would have concluded.

None of this is what happens when I make a photo. Not yet. But it’s aimed at the same gap I live in every time I shoot before I understand what I’m shooting. The gap between having the math and having the eye.

I once asked Doug, a good friend who’s spent as many days on the street as I have, how he knew when to press the shutter. He didn’t have an answer, not really — just a shrug, and something about the moment feeling complete before he could explain why. That shrug took him years to earn. He didn’t keep the years. He kept the shrug.

And then a few years ago Doug did something I still don’t fully understand. He abandoned digital and went back to film. Not for any project, not for the look of it — he could get that in post if he wanted it. He went back to the actual mechanics: loading a roll, metering by hand, often using a tripod, etc. I needled him about it some, the way you’d needle a cigarette smoker who’d taken up a pipe instead, as if the inconvenience were the point. He told me he wanted to slow down, and that film was the only thing that reliably made him do it. Twelve frames and then you stop and reload and you can’t fix it later. The very friction he’d spent decades shooting his way out of, he went looking for again, on purpose.

I don’t know what to do with that, except to notice that he’s the same man who can give me the shrug and also the man who walked back toward the thing the shrug had replaced. Maybe that’s the part the labs haven’t gotten to yet, underneath all the vocabulary of weight updates and meta-learned initializations. Compression is the whole point, until the day it isn’t.

Note: This line of thinking started with a recent essay by Dwarkesh Patel on what he calls continual learning. It’s become a real focus of his thinking about how we get to a better future with AI.

See: https://www.dwarkesh.com/p/the-next-paradigm

Categories
AI

What the Lessor Keeps

Two airlines can fly the same airplane. Not airplanes of the same type — the same airplane, serial number and all, handed back at the end of a lease and reassigned, sometimes within weeks, to a competitor on another continent. AerCap owns more commercial aircraft than any airline on earth, and it leases them to airlines that spend their advertising budgets convincing passengers that flying them is a distinctive experience. The 737 MAX that wears Ryanair’s livery this year might wear Lion Air’s the next, repainted, recertified, its avionics untouched, its airframe indifferent to the change of ownership. The lessor does not care who is flying its asset. It cares that the asset comes back in airworthy condition and that the lease payments clear.

What the airline owns, in the sense that matters, is never the aircraft. It is the route network built up over decades of slot negotiations at constrained airports. It is the maintenance log — every inspection, every part swapped, every anomaly a mechanic in Singapore flagged in 2019 that turned out to predict a fatigue crack nobody else had seen yet. None of that travels with the airplane when the lease ends. It stays behind, compounding, in systems the airline built and the lessor never touches.

Karl Mehta, who has spent a career inside enterprise software watching this kind of asymmetry repeat itself, put a version of it plainly: a model is a brain you rent, and you and your competitor rent the same one. The formulation has the compression of something that has been tested in a few dozen meetings before it found that sentence. It is also, structurally, the airplane story. Anthropic and OpenAI and Google are AerCap. They retain residual value on enormous capital assets — clusters of GPUs depreciating on a schedule, weights trained at a cost that only a handful of balance sheets in the world can absorb — and they lease access to those assets by the token, to anyone who can pay, including, in the same afternoon, two companies trying to put each other out of business. The model does not know whose prompt it is answering. It has no loyalty file. It has, in fact, no memory at all, in the ordinary sense of the word — each call begins exactly where the last one ended for everybody, which is nowhere.

The asymmetry that airlines exploit is the one available here too, and it sits one layer up from the engine. Call it the embedding store, the vector database, the fine-tuning corpus, the retrieval index — the terminology varies by vendor, but the function is constant. It is the accumulated, indexed residue of every customer interaction a company has had, structured so that the rented brain can be handed the relevant fragment of it at the moment of each new call. A bank’s fraud model and a competing bank’s fraud model can call the identical foundation model, route through the identical API, and arrive at entirely different verdicts on the identical transaction, because one of them is retrieving against eleven years of labeled chargebacks specific to its own card portfolio and the other is retrieving against four. The intelligence rented by the hour is, for practical purposes, a commodity, priced down toward marginal cost the way jet fuel is priced — everyone pays close to the same number per unit. The memory is not a commodity. It cannot be, because it is not for sale; it is the institutional record of what has already happened to you, and no amount of capital lets a competitor buy a copy of your chargeback history any more than it lets them buy your maintenance logs.

This produces a particular kind of corporate vertigo, which Mehta’s sentence is really addressing. For three or four years the industry conversation about artificial intelligence has been a conversation about models — which lab’s was larger, which benchmark moved, which release cycle a company should anchor its roadmap to. That conversation rewards being an early and aggressive lessee. But a lessee relationship, however aggressive, does not compound into anything a competitor cannot eventually also lease. The compounding, when it happens, happens in the layer below the API call: in how cleanly a company has structured the record of its own customers, its own failures, its own edge cases, so that the rented brain, plugged in fresh every morning with no memory of yesterday, can be handed exactly the right fragment of yesterday and made to look, for a few hundred milliseconds, like it has been there all along.

A hospital chart has two kinds of entries. There is the vital-signs strip clipped to the bed rail — temperature, pulse, blood pressure, checked every four hours and replaced every four hours, because a reading from yesterday tells the night nurse nothing about the patient in front of her right now. And there is the permanent record in the file downstairs: the allergy that nearly killed him in 2019, the surgery, the medication history going back a decade, written once and never overwritten, because that record is exactly as valuable ten years from now as it is today. Nobody confuses the two charts. Nobody staples last Tuesday’s blood pressure into the permanent file. The hospital figured out, long before anyone digitized it, that memory is not one problem. It is two, and they fail in opposite directions if you run them through the same system.

Most teams building the layer Mehta is describing make exactly that mistake — they staple everything to the same chart. The shorthand for it is dumping everything into a vector database and praying, and it is worth asking why that particular error is so popular. The answer is that it feels like progress: embeddings go in, something resembling memory comes out, and the team moves on to the next sprint without confronting the harder question, which is what kind of memory it just built.

Short-term memory is the vital-signs strip — everything the model needs to finish the task in front of it and nothing it needs after. A customer-service exchange in progress, the order number already mentioned, the fact that this is the second call today, belongs here. So does the scratchpad of a multi-step agent: the search results just pulled, the file just opened, the partial answer being assembled before it commits. The test is not how important the information is but how long it stays true. A customer’s mood this minute is real and gone in twenty minutes; storing it permanently is like stapling yesterday’s temperature reading into the permanent file, undated, until the chart tells you nothing about fever and everything about clutter. Short-term memory should live in the context window itself, or a session-scoped cache, and it should be allowed to die when the session ends. The sin is not forgetting it. The sin is remembering it forever.

Long-term memory is the file downstairs, and it does not come in one shape any more than that file does. The first shape is semantic memory — facts. A customer’s account tier. The chargeback history that decides, in fractions of a second, whether this morning’s transaction clears. Facts belong in a database with a schema, not a vector store, because a fact has a right answer and a vector store gives you an approximate neighbor. Ask a vector index what tier a customer is on and it hands you the five most semantically similar sentences in the corpus — one correct, four merely correct-sounding. Ask a schema the same question and it tells you, because that is what the schema is for.

The more sophisticated shops are already building the seam between the two, rather than picking one and living with its blind spot. A knowledge graph keeps the relationships a schema is good at — this customer, that account, this chargeback, in fixed and queryable connection to one another — while still letting a retrieval layer search across it by meaning rather than by exact key. The approach has a name now, GraphRAG, and the name matters less than what it concedes: that facts and resemblance are different operations, and the honest fix is to run both and let each one answer the kind of question it’s actually suited for, not to force a single index to pretend it can do both jobs at once.

The second shape is episodic memory — what actually happened. The specific conversation last March in which the customer explained, at length, why the previous fix didn’t work. The exact sequence of an agent’s failed attempt at a task, preserved so the next attempt doesn’t repeat it. This is where the vector store finally earns its keep, because an episode isn’t an exact-match lookup, it’s a resemblance — has anything like this come up before — and a vector index, built to find the nearest thing to a fuzzy question, is the right tool for that question and almost no other. The error was never using a vector store. The error is using only a vector store, for facts as well as episodes, on the theory that one hammer with sufficient cosine similarity can stand in for the whole toolbox.

The third shape is the rarest, and the one teams forget to build at all: procedural memory, which is not a fact and not an episode but a skill — the model’s learned sense of how this company writes a refund email, escalates a complaint, formats an invoice. Style is the visible half of it. The other half is harder to see and matters more: the rails the model is forced to run on before it ever gets to choose a word. A refund above some threshold routes to a human, no exceptions, because the workflow says so, not because the model was persuaded to think so on this particular call. An agent that touches a production database does it through a reviewed function with a fixed set of permitted calls, not through whatever query it improvises in the moment. None of that lives in a prompt, and none of it lives in the model’s weights either. It lives in code — the orchestration layer, the permissioning, the state machine the agent is required to pass through — and it is procedural in the oldest sense of the word: not a memory of what to say but a memory of what is and isn’t allowed to happen, enforced whether or not the model that day feels like remembering it. It doesn’t live in a database at all. It lives in fine-tuning, in carefully maintained house-style examples, and in the surrounding scaffolding of guardrails and permitted actions, and it changes slower than the other two, the way a surgeon’s hands carry both technique and caution years after the specific patients are forgotten. A company that has built rich semantic and episodic memory but skipped this layer has a model that knows everything about its customers, writes in exactly the right voice, and is one well-crafted prompt away from doing something the company never agreed to.

The real argument here is not which database serves which layer — that part is plumbing, and plumbing changes every eighteen months. The argument is that memory has to be triaged the way the hospital triages it, with something deciding on purpose what survives the session and what doesn’t, rather than writing every token of every interaction into the same undifferentiated store and trusting retrieval to sort it out later. A vector database with no triage in front of it is not a memory system. It is a landfill with a search function, and it will retrieve the wrong eleven-month-old conversation with the same confidence it retrieves the right one, because nobody wrote the part of the system whose only job is deciding what belongs on which chart.

The lessor’s airplane, repainted, will fly for someone else next year. The route network will not. Neither will the schema that knows a customer’s tier on contact, nor the index that remembers the conversation from last March, nor the fine-tuned hand that knows, without being told twice, how this company writes a refund email. These are the things that do not come back at the end of the lease, because they were never on it.

Categories
AI AI: Large Language Models AI: Transformers Authors Podcasts Writing

The Billboard

The fog was still sitting on the hills when I put in my earbuds and headed out.

Sebastian Mallaby was talking about billboards.

Tim Ferriss had asked him the question he asks everyone: if you could put anything up there, for millions of people to see, what would it be? Mallaby has spent years inside the minds of the people who shaped modern finance — the hedge fund managers, the venture capitalists, the builders of things that changed how the world moves money. He has more material than most people accumulate in a lifetime. He could have said anything.

He said: Prepare your mind.

I kept walking. The houses were quiet in the particular way they get when school lets out for summer — no buses, no car doors, no kids at the corner. Somebody’s sprinklers were running.

The phrase comes originally from Louis Pasteur, who understood something that most people don’t: that chance is not democratic. It does not distribute itself evenly among those who wait. It finds the people who are ready. Chance favors the prepared mind. Pasteur said it, and then he proved it, and then the rest of us spent a century and a half learning it was true.

What struck me about Mallaby’s answer wasn’t the phrase itself. It was the way he said it had kept appearing in his research, surfacing in different decades and different worlds, like a message the material kept trying to send him.

He told the story of Arthur Patterson at Accel Capital. Before a new technology arrived, Accel would work through the implications — what company needs to be built, what founder fits the moment, what the right pitch looks like. So when an entrepreneur finally walked in, when the situation was live and competitive, they already knew ninety percent of what they were hearing. They could move fast because they had already moved slow.

That’s preparation as institutional practice. But Mallaby found the phrase again in a different register entirely, embedded in a single human moment that has always seemed to me like one of the hinge points of our era.

He was interviewing Ilya Sutskever, asking him why he had seen it so quickly.

In 2017, a paper called Attention Is All You Need appeared online. It described a new architecture for neural networks — the transformer — that would eventually rewrite the terms of what artificial intelligence could do. On the day the paper went up, Sutskever read it. And then he ran. He went down the corridor to find his collaborator Alex Radford and told him to stop what he was doing. Everything. Stop. We are going to build a language model on this architecture.

Not someday. Now.

Mallaby asked him how he had seen it so clearly, so fast. And Sutskever’s answer, in its essence, was the same two words: prepared mind.

He had been thinking about the problem of modeling sequential data since his PhD in Canada. For years he had been carrying a question the field hadn’t answered yet. And when the answer appeared — when the transformer showed up on a website one ordinary day — he didn’t have to reason his way toward it. He recognized it. The solution arrived and found a mind that had been waiting for it, that had already cleared space for it, that was already arranged around the shape of exactly this kind of answer.

This is what preparation actually is. Not the accumulation of facts. Not readiness in the generic sense, the vague self-improvement sense. It is the long, patient cultivation of a specific question, held close and kept alive until the answer has somewhere to land.

Mallaby chose that phrase for his billboard because it kept finding him — in the venture capital world, in the AI world, across decades and disciplines and very different kinds of genius. The prepared mind is not a personality trait. It is a practice. It is the work you do before the work arrives.

The sprinklers had clicked off by the time I turned back toward home. The fog was starting to lift off the hills. I was thinking about what I had been preparing for, whether I even knew.

Categories
AI Bicycles History

The Bicycle Shop

Part 2 of 3…

It is eleven-thirty on a Tuesday night and she is arguing with a language model about a spreadsheet.

Not arguing, exactly. That’s not the right word. She is coaxing. She is debugging. She is reading error messages that tell her almost nothing and rewriting prompts that almost work, and she has been doing this for two hours, and the spreadsheet still isn’t right, and she is going to try one more thing before she gives up and does it by hand. She is a data analyst at a mid-sized logistics company in Columbus, Ohio. She is not a researcher. She is not a founder. Nobody is writing about her. She is just a person trying to get a machine to do something useful, and the machine keeps almost doing it, and she keeps learning, in the gap between almost and done, something she couldn’t have learned any other way.

She doesn’t know what she’s learning. That’s the important part.

In 1892, two brothers opened a bicycle repair shop on West Third Street in Dayton, Ohio. The bicycle craze was at its peak — the safety bicycle, with its two equal wheels and chain drive, had just replaced the penny-farthing, that absurd high-wheeler everybody called loose change and the riders, with complete seriousness, called the ordinary. The brothers fixed flats and adjusted brakes and built custom frames and ordered parts from Coventry and kept the books and swept the floor. It was ordinary work. Nobody was writing about them either. What they were doing was accumulating, without knowing they were accumulating, a physical understanding of how machines move through space — the gyroscopic principles, the weight distribution, the thousand small calibrations that kept a rider from falling. They were learning in their hands what no university taught and no book fully contained.

Eleven years later they flew.

We tell the Wright Brothers story as a story about flight. It makes sense — flight is the thing, the miracle, the moment the world changed. But the actual story, the one that explains how Kitty Hawk was possible, is a story about a bicycle shop. It is a story about unglamorous preparatory work, about the education that hides inside the constraint, about what you learn in the gap between the machine that exists and the machine that should exist. Orville and Wilbur didn’t go to Kitty Hawk despite the bicycle shop. They went because of it. The shop was the point. They just didn’t know it yet.

We are in the bicycle shop right now.

The people building with AI today — the prompt engineers, the fine-tuners, the agent builders, the data analysts in Columbus arguing with spreadsheets at midnight — are doing work that looks, from the outside, like mere tinkering. Unglamorous. Iterative. Full of failure. The tools are awkward. The models hallucinate. The context windows run out at the wrong moment. Every solution opens three new problems. It feels like the penny-farthing: powerful enough to be useful, constrained enough to be maddening, requiring a kind of practiced vault just to get started.

But that awkwardness is the education.

Every time a prompt fails, the person writing it learns something about how the model thinks — about what it responds to, what it resists, where it gets confused, where it surprises you. Every agent that breaks in production teaches its builder something about the gap between what a model can do in a demo and what it can do under load, with real data, with users who don’t behave the way you expected. Every context window that runs out forces a decision about what actually matters, what is essential, what can be cut. These are not just technical lessons. They are epistemic ones. They are lessons about the nature of intelligence, about how meaning gets encoded and retrieved, about what it means for a machine to understand something versus to pattern-match on the surface of understanding.

The people learning these lessons right now don’t have a name for what they know. They just know it in their hands.

This is how it always works. James Starley’s craftsmen in Coventry bent and brazed bicycle frames by feel and experience, knowing things in their hands they couldn’t fully explain on paper. That embodied knowledge — the tight tolerances, the interchangeable parts, the discipline of making things that had to work — migrated into every bicycle shop that followed, crossed the Atlantic, and ended up in a shed in Ohio. The Wright Brothers didn’t invent precision manufacturing. They inherited it, absorbed it, and applied it to a problem nobody else had solved because nobody else had brought those particular hands to that particular problem.

The chain drive was the hinge. Before it, the bicycle’s design was locked — bigger wheel for more speed, higher and higher off the ground, until the machine teetered at the edge of what a human could survive. The chain drive broke the constraint. It decoupled the pedals from the wheel, let the gearing do what only size had done before, brought the rider back to earth. What had been a machine for athletes became a machine for everyone. What had been the ordinary became, almost overnight, something new.

We are waiting for the chain drive.

Not waiting passively — it is being built right now, in a hundred places at once, by people who mostly don’t know they’re building it. It might be the interface that finally makes AI genuinely accessible to people who can’t do the running vault. It might be the memory architecture that lets a model carry context the way a human carries context, not in a window but in something more like experience. It might be something nobody has named yet, something that will seem obvious afterward, the way all elegant solutions seem obvious after the fact.

What it will not be is the product of people who stayed away from the bicycle shop.

The analyst in Columbus closes her laptop at midnight. The spreadsheet is still not right. She has learned three things about how the model handles date formatting, two things about how it interprets ambiguous column headers, and one thing about her own assumptions that she didn’t know she was making. Tomorrow she will try again. She will get closer. At some point — not tomorrow, maybe not this year — she will get it right, and the thing she learned in the gap will be available to her for the next problem, and the one after that, and she will carry it forward without knowing she’s carrying it, the way craft always travels, in hands that have done the work.

She doesn’t know what she’s riding toward.

That’s the ordinary part. That’s always been the ordinary part.

Categories
AI Startups

A New Reason to Launch

“Before you launch, the speed you can build is now mainly limited by your imagination in what you tell AI. After you launch, the AI can watch your users and make improvements on its own.”
Jared Friedman, Y Combinator

Jared Friedman watches hundreds of founders a year navigate the gap between idea and launched product. He notices patterns the rest of us miss. And what he’s describing above is not an incremental improvement in how software gets built. It is a change in the nature of the advantage.

This is a different kind of liberation than founders have known before.

The old liberation was launch early and the market corrects your wrong assumptions. Humbling, but useful. You were still the one doing the correcting, late at night, rewriting the onboarding flow based on what the data told you.

The new liberation he’s describing is something closer to multiplication. You launch, and now there are effectively more of you. The AI is watching session replays you’ll never have time to watch. It’s noticing the drop-off after step three that you’d have caught in month four. It’s holding the pattern of a thousand user paths simultaneously and asking what they mean. Your imagination seeded the thing. Reality is now feeding it.

That observation redraws the map cleanly. Pre-launch and post-launch used to differ in degree — you knew more after than before. Now they differ in kind. Pre-launch you are the sensing organ. Post-launch you’ve grown new ones.

The founders who feel this most viscerally, I suspect, are the ones building alone or in pairs — the people for whom every previous era of building had a hard ceiling imposed by human hours. They could only read so many support tickets. They could only run so many experiments. The ceiling is lifting and the feeling is of a room getting larger.

The core advice hasn’t changed. Paul Graham was saying “launch early” twenty years ago and it was true then. What’s changed is the reason underneath it — the mechanism that makes it true now is nothing like the one he had in mind.

The advice is twenty years old. There is a new reason and it is brand new. Most people haven’t noticed the swap yet. But they will.

That window does not stay open long.