Categories
AI

The Encyclopedia and the Reasoner

I was standing in the cereal aisle a few weeks ago, doing the thing I always do โ€” flipping the box over, scanning the fine print, comparing fiber grams like it mattered more than it probably does โ€” when I thought about the model I’d been testing that morning. Sharp. Fast. Occasionally, confidently, wrong about something I could have looked up in ten seconds.

There was no label for that. No panel telling me what was inside, what it was good at, what it might get wrong, what it cost to run. Just a chat window and a kind of blind trust.

That’s the itch behind this post. What would it look like if AI models came with something like a Nutrition Facts label โ€” the kind the FDA forced onto every box in your pantry back in 1994? Not as a gimmick, but as a real answer to a real problem: we are feeding these things into our decisions, our writing, our portfolios, our kids’ homework, largely on faith.

The IQ Number That Isn’t Quite an IQ Number

I keep running into a shorthand in investing circles โ€” Jordi Visser and others talking about frontier models as “140 IQ” systems, reasoning at a level that outpaces most humans on the kinds of puzzles we associate with fluid intelligence. Pattern recognition. Logic chains. Novel deduction under pressure.

It’s a useful number. It’s also a bit of a trick.

Human IQ tests were built to measure something narrow and specific โ€” not wisdom, not knowledge, not judgment, but the raw machinery of reasoning. When we borrow that language for AI, we inherit the same narrowness, which is fine as long as we remember it. A model that aces abstract reasoning benchmarks isn’t necessarily the model that knows the correct dosage, the right case law, or what actually happened in 1932. Reasoning and knowledge are cousins, not twins.

Two Kinds of Smart

Here’s an old-fashioned way to think about the split: Britannica versus World Book.

Britannica was the encyclopedia my father would have trusted โ€” dense, expert-written, unapologetically deep, assuming you could keep up. World Book was the one actually sitting on the shelf in most houses I knew growing up, mine included: friendlier, broader, built for a general reader, a little shallower in exchange for being a little more useful on a Tuesday night with a homework assignment due.

Neither is wrong. They’re optimized for different things. And training data does the same kind of sorting. A model fed heavily on curated, scholarly, expert-vetted sources leans Britannica โ€” deep, careful, occasionally slow to update. A model trained on the sprawl of the open web leans World Book โ€” broad, current, occasionally sloppy, sometimes brilliant at the edges precisely because it’s seen everything.

Any honest label for a model needs a section on this. Call it “Knowledge Sourcing.” Not just how big the training set was, but what kind of encyclopedia it’s pretending to be.

Sketching the Label

If I could design the box myself, it might read something like this:

Serving Size: 1 query, ~500 tokens

Reasoning Score: 138 (fluid problem-solving, logic, abstraction) Knowledge Depth: Moderateโ€“High (cutoff: [date]; strongest in [domains]; weakest in [domains])
Ingredients: Curated scholarly corpora, licensed news archives, public web crawl, synthetic reasoning data, human feedback Allergens: Confident hallucination under ambiguous prompts; recency gaps beyond training cutoff; known weakness in [specific domain]
Cost per Serving: $X per million tokens; Y watt-hours per query Best Paired With: Retrieval tools, human review for high-stakes decisions

It’s a little tongue-in-cheek written out like that. But underneath the joke is something I actually want โ€” the same instinct that made me read cereal boxes as a kid. Not to be scared of what’s inside, just to know.

The Part That Actually Excites Me

Here’s where the scaling laws get interesting, and where I think the real opportunity sits.

World knowledge is expensive. It’s greedy for data and parameters โ€” you need to have practically read the internet to know the boiling point of tungsten, the plot of a minor Victorian novel, and the org chart of a mid-cap company all at once. Reasoning, it turns out, is a different kind of animal. It can be distilled, compressed, taught through synthetic problems and careful post-training, and squeezed into something far smaller than you’d expect.

Which means a genuinely thrilling possibility is already taking shape: sharp, high-reasoning models small enough to run on a phone or a laptop, entirely offline, because they’ve shed the encyclopedia and kept the mind. Pair one of those with a personal index โ€” your own notes, your own documents, a retrieval layer built around your actual life โ€” and you get something closer to a personal thinking partner than a general-purpose oracle. Private. Fast. Always available. Tuned to you rather than to everyone. Apple may be on to something with this kind of strategy?

I think about this constantly in my own workflow โ€” the daily scans, the little agents I’ve built to help sort signal from noise, the genealogy digging, the investment frameworks I keep refining. What I usually want isn’t more encyclopedia. It’s a clear-headed reasoner sitting next to my own carefully kept knowledge, not buried under someone else’s version of the whole internet.

Why the Label Matters More Than the Score

None of this works, though, without honesty about what’s inside the box. A 140 on a reasoning benchmark tells you almost nothing about whether a model will quietly misremember a fact it was never that confident about in the first place. And a model can be extraordinarily knowledgeable while being a mediocre reasoner โ€” plenty capable of reciting the right ingredients and still getting the recipe wrong.

The nutrition label movement in food didn’t eliminate junk food. It just made it possible to choose junk food on purpose, with your eyes open, instead of by accident. I’d like the same deal with AI. Not a demand that every model be a genius generalist, but a demand that I get to know what I’m actually consuming โ€” and choose the lean local thinker over the bloated encyclopedia when that’s what the moment calls for, or the other way around when it isn’t.

Curiosity got me into that cereal aisle habit decades ago, and it’s the same instinct pulling me toward this idea now โ€” not suspicion of the box, just a wish to read it clearly before I decide how much of it to trust.

What would you want on your label?

Categories
AI Business Investing Technology

The Scarcity Portfolio: Navigating Sovereign Debt, Wafer Bottlenecks, and Orbital Compute

Today I was watching the interview of Gavin Baker by Patrick Oโ€™Shaughnessy on his Invest Like the Best podcast. Like prior conversations this was another fascinating excursion into the mind of a sophisticated and very successful tech venture investor.

During the conversation, Patrick asked Gavin what agents he was using that were especially helpful and he mentioned one which summarizes YouTube podcasts and videos for him. Like most of us Baker just doesnโ€™t have the time to watch or listen to them himself so good summaries are really helpful.

Turns out Iโ€™ve been working on a Google Gemini Gem that does this for me. When Baker mentioned his I fired up the new Gemini 3.5 Flash model and asked it to summarize the Baker interview.

Later in the conversation Baker used the term โ€œbattlefield AIโ€ which caused me to go back to Gemini again to learn more about that. The results were so interesting that I asked Gemini to create a syllabus for a semester class on these subjects. After that I asked it to convert our whole conversation into a Markdown file so I could share it. Youโ€™ll find it below.

I found this whole experience pretty stunning. I came away very impressed with Gemini 3.5 Flash both for the quality of the responses but also the sheer speed. Wow!

Anyway I hope you enjoy the following!


Categories
AI Anthropic Business Google

The Weight of the Bill

Jordi Visser has been making the case for months โ€” in his weekly YouTube commentary and on his Substack โ€” that we are living through an exponential transition that most people are measuring with the wrong instruments. I think he’s right. I found two data points this week that suggest why.

I was somewhere in the middle of an Invest Like the Best episode when Dylan Patel said it โ€” almost as an aside, the kind of thing you drop to establish context before moving on to the point you actually came to make. His firm, SemiAnalysis, analyzes the semiconductor and AI industries for a living. And their usage of Claude, he noted, has been growing. The costs have been growing too.

Exponentially.

He moved on. I didn’t.

I think Patel’s API bill might be one of the more honest documents in the current AI moment โ€” more honest than the analyst reports his firm produces, more honest than the earnings calls where every public company performs its AI fluency for shareholders.

Surveys bend. When you ask someone whether they’re using AI in their work, you’re asking them to self-report on a technology that has become a proxy for relevance, for not being left behind. The incentive to say yes is enormous. And even when the yes is genuine, it tells you nothing about depth โ€” whether AI has become load-bearing in how someone actually works, or whether it’s an impressive thing they do occasionally.

Nobody pays exponentially growing API costs for show. Money is the honest witness.

What makes Patel’s situation quietly strange is the recursion in it. SemiAnalysis exists to help sophisticated investors and technologists understand this industry โ€” and they cannot predict their own consumption curve. They are inside the exponential the same way everyone else is. They just happen to be watching their bill.

Then this morning, a different number arrived. Google announced it will invest up to $40 billion in Anthropic โ€” $10 billion committed now, another $30 billion contingent on performance milestones. This follows a separate $5 billion from Amazon, part of a broader arrangement under which Anthropic is expected to spend up to $100 billion on compute over time.

The temptation with numbers like these is to treat them as spectacle. Forty billion dollars is so large it becomes almost aesthetic โ€” a statement about ambition, about the kind of bets that define eras. You feel the weight of the zeros and move on.

But I keep coming back to Patel’s API bill.

Because Google’s $40 billion and SemiAnalysis’s compounding monthly costs are saying the same thing, expressed at scales so different they almost don’t seem related. One is a research firm noticing that their tool usage has quietly escaped prediction. The other is one of the most sophisticated capital allocators on earth making a bet that strains comprehension. But both are pointing at the same reality: that this technology, wherever it takes hold, does not plateau. It compounds.

We have been waiting, I think, for the moment when AI adoption becomes legibly real โ€” some threshold event that separates the signal from the noise, the press release from the actual change. The surveys were supposed to mark that moment. The enterprise announcements. The benchmark numbers.

Patel’s aside suggests we’ve been waiting for the wrong thing. You don’t arrive at the exponential. You just eventually notice you’re already in it โ€” in an aside on a podcast, before moving on to the point you actually came to make.