Categories
AI

The Encyclopedia and the Reasoner

I was standing in the cereal aisle a few weeks ago, doing the thing I always do — flipping the box over, scanning the fine print, comparing fiber grams like it mattered more than it probably does — when I thought about the model I’d been testing that morning. Sharp. Fast. Occasionally, confidently, wrong about something I could have looked up in ten seconds.

There was no label for that. No panel telling me what was inside, what it was good at, what it might get wrong, what it cost to run. Just a chat window and a kind of blind trust.

That’s the itch behind this post. What would it look like if AI models came with something like a Nutrition Facts label — the kind the FDA forced onto every box in your pantry back in 1994? Not as a gimmick, but as a real answer to a real problem: we are feeding these things into our decisions, our writing, our portfolios, our kids’ homework, largely on faith.

The IQ Number That Isn’t Quite an IQ Number

I keep running into a shorthand in investing circles — Jordi Visser and others talking about frontier models as “140 IQ” systems, reasoning at a level that outpaces most humans on the kinds of puzzles we associate with fluid intelligence. Pattern recognition. Logic chains. Novel deduction under pressure.

It’s a useful number. It’s also a bit of a trick.

Human IQ tests were built to measure something narrow and specific — not wisdom, not knowledge, not judgment, but the raw machinery of reasoning. When we borrow that language for AI, we inherit the same narrowness, which is fine as long as we remember it. A model that aces abstract reasoning benchmarks isn’t necessarily the model that knows the correct dosage, the right case law, or what actually happened in 1932. Reasoning and knowledge are cousins, not twins.

Two Kinds of Smart

Here’s an old-fashioned way to think about the split: Britannica versus World Book.

Britannica was the encyclopedia my father would have trusted — dense, expert-written, unapologetically deep, assuming you could keep up. World Book was the one actually sitting on the shelf in most houses I knew growing up, mine included: friendlier, broader, built for a general reader, a little shallower in exchange for being a little more useful on a Tuesday night with a homework assignment due.

Neither is wrong. They’re optimized for different things. And training data does the same kind of sorting. A model fed heavily on curated, scholarly, expert-vetted sources leans Britannica — deep, careful, occasionally slow to update. A model trained on the sprawl of the open web leans World Book — broad, current, occasionally sloppy, sometimes brilliant at the edges precisely because it’s seen everything.

Any honest label for a model needs a section on this. Call it “Knowledge Sourcing.” Not just how big the training set was, but what kind of encyclopedia it’s pretending to be.

Sketching the Label

If I could design the box myself, it might read something like this:

Serving Size: 1 query, ~500 tokens

Reasoning Score: 138 (fluid problem-solving, logic, abstraction) Knowledge Depth: Moderate–High (cutoff: [date]; strongest in [domains]; weakest in [domains])
Ingredients: Curated scholarly corpora, licensed news archives, public web crawl, synthetic reasoning data, human feedback Allergens: Confident hallucination under ambiguous prompts; recency gaps beyond training cutoff; known weakness in [specific domain]
Cost per Serving: $X per million tokens; Y watt-hours per query Best Paired With: Retrieval tools, human review for high-stakes decisions

It’s a little tongue-in-cheek written out like that. But underneath the joke is something I actually want — the same instinct that made me read cereal boxes as a kid. Not to be scared of what’s inside, just to know.

The Part That Actually Excites Me

Here’s where the scaling laws get interesting, and where I think the real opportunity sits.

World knowledge is expensive. It’s greedy for data and parameters — you need to have practically read the internet to know the boiling point of tungsten, the plot of a minor Victorian novel, and the org chart of a mid-cap company all at once. Reasoning, it turns out, is a different kind of animal. It can be distilled, compressed, taught through synthetic problems and careful post-training, and squeezed into something far smaller than you’d expect.

Which means a genuinely thrilling possibility is already taking shape: sharp, high-reasoning models small enough to run on a phone or a laptop, entirely offline, because they’ve shed the encyclopedia and kept the mind. Pair one of those with a personal index — your own notes, your own documents, a retrieval layer built around your actual life — and you get something closer to a personal thinking partner than a general-purpose oracle. Private. Fast. Always available. Tuned to you rather than to everyone. Apple may be on to something with this kind of strategy?

I think about this constantly in my own workflow — the daily scans, the little agents I’ve built to help sort signal from noise, the genealogy digging, the investment frameworks I keep refining. What I usually want isn’t more encyclopedia. It’s a clear-headed reasoner sitting next to my own carefully kept knowledge, not buried under someone else’s version of the whole internet.

Why the Label Matters More Than the Score

None of this works, though, without honesty about what’s inside the box. A 140 on a reasoning benchmark tells you almost nothing about whether a model will quietly misremember a fact it was never that confident about in the first place. And a model can be extraordinarily knowledgeable while being a mediocre reasoner — plenty capable of reciting the right ingredients and still getting the recipe wrong.

The nutrition label movement in food didn’t eliminate junk food. It just made it possible to choose junk food on purpose, with your eyes open, instead of by accident. I’d like the same deal with AI. Not a demand that every model be a genius generalist, but a demand that I get to know what I’m actually consuming — and choose the lean local thinker over the bloated encyclopedia when that’s what the moment calls for, or the other way around when it isn’t.

Curiosity got me into that cereal aisle habit decades ago, and it’s the same instinct pulling me toward this idea now — not suspicion of the box, just a wish to read it clearly before I decide how much of it to trust.

What would you want on your label?

Categories
AI Photography

The Price of the Cold

Two men are standing close to a brick wall trying not to talk, because talking wastes what little warmth is left in a body that has been outside too long. One of them has a camera — Jerry Schatzberg, a fashion photographer. His hands are jammed half into his coat pockets between shots. The other man has his collar up around his ears and a scarf wound twice, black and white, and he is not moving much, because moving costs heat, and heat is the one thing neither of them has enough of. Schatzberg raises the camera. His fingers, by this point, are not entirely his own. When he presses the shutter there is a tremor in it he did not order and cannot undo.

The picture comes out smeared at the edges. Bob Dylan’s face, in the frame, is dissolving slightly into the gray behind him, like a man photographed through a windshield in the rain. It is, by any studio standard, a bad photograph. Schatzberg knows it’s a bad photograph. He has made a career out of not taking bad photographs.

And it became the cover of Blonde on Blonde, which is the best rock album ever recorded, and in nearly sixty years nobody has managed to improve on it by reshooting it clean. The blur isn’t a decision. It’s a symptom — of two men standing in the cold too long, of a photographer choosing, afterward, to keep the evidence of his own discomfort instead of erasing it.

There’s a difference between an accident and serendipity that I don’t think gets said out loud enough, and it matters more than it used to. An accident is the cold — involuntary, uninvited, spent before you know if it was worth spending. Schatzberg didn’t choose to shiver. His hands moved because his body was doing what bodies do at a certain temperature, and the shutter caught what his hands actually did, not what he meant to do. Serendipity is what happens next: a verdict, rendered after the fact, that the wreckage of an intention was better than the intention itself. The accident is what makes the verdict possible. Without the cold, there’s nothing to render a verdict on.

I’ve been sitting with a large language model most days for the better part of a year now, watching it write, asking it to try again, watching it try again in a way that is never quite the same and never quite different enough to matter. Somewhere upstream of me there is a number called temperature, and I will never see it. Somebody else did, once, in a meeting, and decided that the word for controlled, pre-approved, refundable randomness should be temperature — the same word for the thing that made Schatzberg’s hands shake, the same word for the actual physical stakes of standing outside too long in January without enough coat — and then set it, and moved on, and nobody in that meeting laughed, because nobody in the room had ever been cold in a way that mattered to the work.

Picture the room instead. It is climate-controlled to sixty-eight degrees, humidity held flat, year-round, by a building management system nobody thinks about until it fails. Somewhere in it, the hardware is generating your next five versions of a photograph like the one on Blonde on Blonde. Nobody in that room is going to lose feeling in their fingers today. Nobody’s collar is up. I don’t know his name — nobody outside the building does — but somebody like him tuned the sampling distribution and went home at six. That’s the guy in the good suit. He built the weather. He never once stood in it.

The small model inherits conclusions. It never inherits the cold. Whatever accidents shaped the teacher model’s own training — whatever costly friction produced the insight in the first place — the student model gets none of that weather. It gets the photograph, cropped and sharpened, with the blur removed because somebody along the way decided the blur was noise instead of signal — the way Schatzberg, a lesser photographer, might have reshot Dylan clean and thrown the bad one away. It is heir to a serendipity it never earned, because it was never present for the accident that made the serendipity possible. It is, in the most literal sense the industry means by the word, cheap.

I keep coming back to the fact that nobody at the API layer is shivering. That’s not a complaint, exactly. It’s just an observation about where the cost went. Somewhere in the training data, some human being was cold, or scared, or holding a fish that was starting to smell, or standing on a stepladder with ten minutes before the traffic came back, and that person paid a real price for a result they couldn’t yet know was good. The model downstream of all that gets the result without the price.

Two rooms, then. In one of them it is January in New York and a man’s fingers have stopped entirely obeying him. In the other it is sixty-eight degrees, always, on a Tuesday and on a Sunday and at three in the morning, and the machines are making you nine more versions of that same blur. Sixty-eight degrees. A number, upstream, that you will never see.