Categories
AI AI: Prompting

The Price of the Barstool

Some genius out in California โ€” one of the AI guys, the smart ones, always the smart ones โ€” says the trick to getting a machine to understand you is to stop typing and start talking. Ramble for ten minutes, he says. Total mess. Say whatever’s in your head, contradict yourself, circle back, don’t clean it up. The machine, he says, is better at finding what you meant than you are.

I could’ve told him that thirty years ago for the price of a cup of coffee, except I would’ve told him to go find a bartender.

Every good bartender in this city has been doing this since before anybody could spell computer. You sit down, you’re a mess, you talk for twenty minutes about your ex-wife and your kid’s tuition and the guy at work who’s getting the promotion you should’ve gotten, and it comes out in no particular order, and somewhere in the fourth minute you say the one true thing โ€” I think I’m afraid I already peaked โ€” and the bartender doesn’t say a word, just wipes the bar down in front of you, and by the time you leave you know something about yourself you didn’t know when you walked in. Nobody paid Sam a nickel for that. Nobody gave him a paper to publish.

Now they’ve built a machine that does the same trick and they’re calling it a pattern. They put it in a paper. Somebody in Palo Alto’s going to raise money on it.

Here’s what I’ll say for the guy โ€” he’s right, and being right is rarer than people think in that business. The mess is where the truth lives. Nobody ever said something true the first time they tried to say it cleanly. You clean it up too fast, you’ve written a memo. You let it run messy, you’ve said something.

But I’ll tell you what he won’t put in the paper, because it doesn’t fit on a slide. The machine will hand you back a cleaner version of your ten minutes, and it’ll sound better than what you said, and half the time it’ll be missing the one part that was actually true, because the true part usually comes out sounding wrong the first time. That’s the part a bartender knows and a machine doesn’t. Sam wouldn’t have cleaned up your sentence. He’d have just remembered you said it, and brought you another one, and let you sit there with it.

They’ll charge you a subscription for the part where it listens. The part where somebody sits with what you actually said โ€” that’s still free, if you can find the right stool.

Categories
Computers IBM

The Day the Last Mainframe Went Dark

Note: I literally grew up during the heyday of the IBM mainframe era. My first real job was working for IBM in San Francisco beginning in 1968. Iโ€™m a โ€œbig ironโ€ kind of guy. But this post was imagined after reading the following in a July 14, 2026 announcement from IBM: When we discussed our expectations with you in April, we noted that we would be wrapping on the launch of z17 in the second quarter. Given this was the strongest start to a mainframe program in our history, we expected Infrastructure revenue to decline low-single digits for the year, beginning this quarter. What played out was worse than our expectations, driven by a shortfall in our Z performance and the associated software stack, primarily in Transaction Processing. In the last few weeks of June, we saw clients shift their quarterly capex spend toward servers, storage, and memory purchases to secure supply-constrained infrastructure ahead of expected price increases. This dynamic impacted client buying patterns. While we anticipated some supply chain related impact in our expectations, we did not anticipate the magnitude of the capex reprioritization.


It won’t arrive with fanfare. No countdown, no viral video of engineers raising a glass. One morningโ€”sometime in the 2040s, perhaps laterโ€”a small team in a climate-controlled data center will complete the final cutover. They’ll flip the switches, watch the lights dim, and listen as the hum of the last production IBM mainframe fades to silence. An era that began with the System/360 in 1964 will end. Not with a crash. With the soft click of obsolescence.

We’ve been predicting the mainframe’s death for decades. In the early 1990s, pundits declared it doomed. They were wrong. Those systemsโ€”reliable, secure, capable of staggering transaction volumes with near-perfect uptimeโ€”became the invisible backbone of modern life. Your last bank transfer, airline reservation, insurance claim, or government benefit likely touched one. They endured because they solved hard problems well: high-volume, mission-critical processing where failure was never an option.

The path to that final power-down was never a rupture. It was a long, uneven evolutionโ€”driven by economics, technology, talent shifts, and the patient work of modernization. AI tools accelerated the transition. They didn’t cause it.

Lessons from Earlier Transitions

Steam engines dominated railroads for generationsโ€”powerful, reliable, deeply integrated into the industrial economy. Diesel won through incremental advantages: better efficiency, lower maintenance, longer trains with less labor. Railroads rebuilt infrastructure and retired the old iron as the economics aligned, route by route.

Prop planes opened the skies to mass travel. Jets brought speed and range that transformed global connectivityโ€”but airlines didn’t scrap fleets overnight. They ran hybrids during the overlap, invested in new airports and training, and retired props as jet economics and passenger demand made the case irresistible.

The mainframe followed this pattern. AI coding agentsโ€”Claude Code, Cursor, OpenAI models, AWS Transform, IBM watsonxโ€”transformed the brutal manual work of understanding undocumented COBOL, extracting buried business logic, generating tests, refactoring safely. What once demanded scarce veteran experts for months or years could now be accelerated, with rigorous human oversight and equivalence testing.

Platforms like Visa’s Pismo showed a smarter path: incremental modernization. Cloud-native microservices layered alongside legacy cores, rather than rip-and-replace. Banks demonstrated real progress. Hybrid strategies wonโ€”AI inference running close to sensitive data on evolved mainframes (IBM’s z17 and successors, with on-chip accelerators), while new applications and analytics moved to elastic cloud environments.

IBM positioned the platform as an “AI factory” for low-latency, secure workloads. But pricing pressure was constant. High, capacity-based software licensing made the economics harder to defend as cloud offered predictable, usage-driven costs and younger talent gravitated toward modern stacks. For CFOs weighing rising maintenance against retiring COBOL expertise and AI-assisted migration, the scales tipped.

By the mid-2030s, competitive and regulatory forces intensified. Fujitsu’s exit from mainframes created a cliff in affected markets. Skills shortages accelerated. Even the most conservative holdoutsโ€”ultra-high-volume, regulated systems in finance, government, specialized industriesโ€”began serious moves, as simulation environments and exhaustive parallel testing brought the risk down to manageable size.

The Final Act

The last systems to go were the stubborn ones, where disruption carried outsized consequences. When the final cutover succeededโ€”after months of flawless parallel runningโ€”the team powered down the machine. A global bank, a payments processor, a government entity. Maybe a small ceremony: engineers who’d kept it alive for decades, trading stories of the iron that never failed when the world needed it most.

Picture the aircraft boneyards outside Tucson or Victorville, retired 747s sitting in rows under the sun, giving up parts to new generations before they’re recycled. Mainframes will meet a similar fate, more climate-controlled. Some linger in warehouses as insurance, still humming faintly in test or archival roles. Others get dismantled by IT asset disposition teamsโ€”data wiped to standard, processors and I/O cards harvested for niche markets. The bulk gets recycled, metal and circuitry returning to the supply chain. Like the jets, the iron won’t vanish in disgrace. Its lessons in reliability and disciplined engineering at scale live on, embedded in whatever comes next.

The world didn’t stop. Transactions kept flowing, now on distributed, elastic, AI-augmented platforms that had absorbed the best of what came before. The mainframe era didn’t end in failure. It ended because better options finally existed for every workload.

What Endures

We’ll look back with respect and nostalgia. The mainframe wasn’t flashy, but it taught something durable: some problems reward obsessive focus on reliability and scale; disciplined engineering outlasts hype cycles; the wisest transitions are rarely clean breaks. They’re patient evolutions that carry forward what matters.

IBM will have completed its own transformation by thenโ€”software, services, hybrid orchestration, AI tools that work across environments. The company that built the platform helps close the book on it.

The last mainframe going dark won’t feel like loss. It will feel like the natural close of a chapter that powered the digital economy through its most formative decades. The iron did its job. Now the next architecture takes the stage, standing on shoulders built to last.

Categories
AI China

Model Competition

The Chinese AI company Moonshot AI released its Kimi K3 model yesterday. I played with it a bit in the Kimi iOS app and, frankly, was underwhelmed. It was slow and disappointing. But this was hardly an extensive test. Itโ€™s strictly my first impression. Iโ€™m not doing coding. Iโ€™m what model makers would call a knowledge worker.

This morning Gavin Baker posted on X about Kimi K3 calling it potentially an โ€œimportant inflection point for AIโ€:

Kimi K3 may be an important inflection point for AI. Potentially negative for Anthropic and OpenAI while being net positive for essentially every other company in the world. I mean that very literally. Although the real โ€œSputnik momentโ€ would be an open-source frontier model that was also token efficient unlike Kimi K3 which is 50-70% more expensive to run than GPT 5.6 per Artificial Analysis.

Rationale:
 
A world where there are only 2-3 dominant frontier labs with 90% inference margins is net negative for every other layer while being awesome for those 2-3 labs. Those labs would become monopsonies for power, data centers, semiconductors and hyperscalers and would obviously vertically integrate over time into all those layers while also completely subsuming the application/software layers. 
 
Anything that lowers margins and increases competition at the model layer is good for every other AI layer: power, semiconductors, hyperscalers, neoclouds and yes even software.
 
This is why Jensen is so supportive of open-source. An open-source model requires the exact same amount of compute to run as a closed frontier model of similar size and architecture. Kimi K3 is roughly the same price as GPT 5.6 Terra on a per token basis, which actually suggests that it is less computationally efficient as I am sure that GPT 5.6 is priced to a higher margin than K3. And given that K3 is a token wastrel, i.e. token inefficient, it is significantly more expensive per task than GPT 5.6 and Grok 4.5, which are much more token efficient. Cost per token and token efficiency (i.e. intelligence density per token) are the drivers of intelligence per unit of cost. The winning AI companies will be those that offer the most intelligence per $ over time.
 
Lower margin % at the model layer = more margin $ at every part of the infrastructure layer and is a godsend for software. This can happen either through open-source models like K3 at the frontier or having a vertically integrated model company like Meta, SpaceX or Google at the frontier. Both outcomes result in a lower margin % at the model layer as vertically integrated model companies donโ€™t really care where the margin $ come from. This is why it was so painful for OpenAI and Anthropic when Google was right there with them from a model competitiveness perspective and why Grok 4.5 and Muse 1.1 were just as important as Kimi K3.
โ€จThe reason Kimi K3 is only potentially negative for Anthropic and OpenAI is 1) the @ericvishria point that the Claude and ChatGPT products and harnesses may be more important than their models today and 2) the hypothesis that they have much more advanced model checkpoints internally that are already being used for RSI. In the latter scenario, reaching RSI even a few months ahead of other labs might be enough to cement a permanent lead.

Time will tell on both points. And likely fairly quickly.

Caveat would be that since Kimi K3 is not token efficient and thereby actually more expensive than ChatGPT 5.6, we may need to see a more token efficient open-source model at the frontier or see Grok 5/Composer 4/Muse 2 at multiple points on the Pareto frontier for this potential risk to Anthropic and OpenAI to play out. And I am sure they will both vertically integrate as quickly as possible while continuing the product/harness strength they have shown over the last 8 months.

I wanted to explore this further so I asked Grok to explain his post:

This is a tweet (X post) from Gavin Baker (@GavinSBaker), Managing Partner and CIO of Atreides Management, posted on July 17, 2026.

It analyzes the recent launch of Kimi K3 (by Chinese AI company Moonshot AI) as a potentially significant shift in the AI industry. Baker argues it could hurt the high-margin dominance of closed U.S. frontier labs like OpenAI and Anthropic while benefiting nearly everyone else in the broader AI ecosystem (infrastructure, chips, data centers, power, hyperscalers, and software).

Quick Context on Kimi K3

Kimi K3 is Moonshot AIโ€™s new flagship model (launched ~July 16, 2026). Key specs include:

  • ~2.8 trillion parameter Mixture-of-Experts (MoE) architecture.
  • Massive 1 million token context window.
  • Native vision capabilities.
  • Strong early performance in coding, agentic tasks, and long-context reasoning (competitive with or close to top closed models like GPT-5.6 variants or Claude models in certain benchmarks).
  • Pricing: $3 input / $15 output per million tokens (with caching discounts).
  • Important: Full weights are scheduled for open release (modified MIT license) around July 27, 2026 โ€” making it one of the first truly open frontier-scale models.

It represents a step up from prior Kimi models and adds meaningful competition from a non-U.S. lab.1

The Two Images in the Tweet

These are screenshots from Artificial Analysis (an independent benchmarking site):

  1. Bar chart โ€” โ€œCost per Intelligence Index Taskโ€ (lower is better). It breaks down costs by token type (input, output, reasoning, cache, etc.) and shows Kimi K3 as relatively expensive compared to some GPT-5.6 variants.
  2. Scatter plot โ€” Intelligence Index (y-axis) vs. Cost per task (x-axis, log scale). It positions models in quadrants, with the โ€œmost attractiveโ€ area being high intelligence at low cost. Kimi K3 appears competitive on intelligence but not the cheapest.

These visuals support Bakerโ€™s point that while Kimi K3 reaches frontier-level intelligence, it is not the most token-efficient, making its effective cost per useful output higher than more efficient models.

Bakerโ€™s Core Argument (Broken Down)

Bakerโ€™s thesis is economic and structural, not just technical:

  • A world with only 2โ€“3 dominant closed frontier labs (with ~90% inference margins) is bad for almost everyone else.
    Those labs would act as monopsonies (dominant buyers) for power, chips, data centers, and compute. They would vertically integrate upward (into infrastructure) and downward (into apps/software), capturing most of the value. Competition at the model layer prevents this concentration.
  • Anything that increases competition or lowers margins at the model layer is broadly positive.
    It redistributes economic value to the rest of the stack: more demand for GPUs/chips (inference especially), power, data centers, hyperscalers/neoclouds, and software (because cheaper/better AI drives wider adoption).
  • Token efficiency matters more than raw price per token.
    Baker notes Kimi K3 is roughly the same per-token price as GPT-5.6 but 50-70% more expensive overall according to Artificial Analysis. It is a โ€œtoken wastrelโ€ โ€” it uses more tokens to achieve similar results (lower โ€œintelligence density per tokenโ€).
    โ†’ True winners deliver the most intelligence per dollar. Cost-per-token + efficiency together determine real value.
  • Why NVIDIAโ€™s Jensen Huang loves open-source models.
    Open models still require massive compute to run at scale (same as closed ones). More models + wider adoption = more inference demand = more GPUs sold. A few ultra-high-margin closed labs would concentrate spend and reduce overall volume.
  • Vertical integration changes the game.
    Companies like Google, Meta, or xAI/SpaceX that have both strong models and their own infrastructure donโ€™t obsess over model-layer margins โ€” they capture value elsewhere. This is why more competition (open or vertically integrated) hurts pure-play high-margin model companies like OpenAI/Anthropic.
Why Itโ€™s Only โ€œPotentiallyโ€ Negative for OpenAI and Anthropic

Baker is careful:

  • Their products, user interfaces, and โ€œharnessesโ€ (the surrounding tools/ecosystems) may matter more than raw model performance right now.
  • They likely have more advanced internal models/checkpoints (possibly already used for recursive self-improvement/RSI).
  • Kimi K3โ€™s token inefficiency softens the blow for now. A more efficient open-source frontier model (or stronger showings from Grok, etc.) would be more threatening.
  • OpenAI and Anthropic are already moving to vertically integrate.
Overall Takeaway from the Post

Baker sees Kimi K3 (and the wave of strong Chinese models like Qwen, DeepSeek, GLM) as healthy competition that prevents unhealthy concentration. Itโ€™s not yet a full โ€œSputnik momentโ€ (a truly dominant, efficient open-source frontier model), but it points in that direction.

Net effect he describes: Lower % margins at the model layer โ†’ more total margin dollars flowing to infrastructure, chips, power, and software layers. This is bullish for the broader AI ecosystem even if it pressures the pure closed frontier labs.

The tone is analytical and balanced โ€” he acknowledges uncertainties and that time will tell (likely quickly). Replies to the post generally agree this kind of competition is positive for the industry overall and could shift more spend toward efficient inference hardware.

This fits Bakerโ€™s broader investing views on AI: he often emphasizes infrastructure winners, token economics, and how value accrues across the full stack rather than just at the smartest model.

Interesting.

Categories
AI Business

The Reverse Information Paradox We’ve Always Had

Satya Nadella wrote recently about what he calls the Reverse Information Paradox: enterprises pay for AI intelligence twice. Once in money. Again in the proprietary knowledge they surrender through every prompt, correction, and evaluation. The better they use the model, the more of their own institutional understanding leaks into someone else’s system. The vendor ends up knowing more about the buyer’s business than the buyer knows about what the vendor retained.

Replace “model” with “employee” (or โ€œconsultantโ€) and the paradox is not new at all.

You pay for a person once with salary. You pay again with something harder to price: the context, relationships, and judgment they must absorb to become useful to you. The better they perform, the deeper the immersion, the more of your particular way of doing things moves into their head. Every correction and late-night conversation is another trace of institutional memory changing hands. When they leave, some of that memory leaves with them. Not always through theft. Usually just through the ordinary residue of good work.

The visible cost is salary; the invisible cost is the slow transfer of what makes you distinctive. High performers get more access precisely because they’re high performers, which means the leakage accelerates exactly when you can least afford it. The exhaust is just harder to see with people than with tokens โ€” it moves through conversation and mental models instead of logs.

The analogy has a limit, and the limit matters. Employees bring knowledge in, not just absorb it. They have judgment and relationships a model doesn’t. Models are purely absorptive, and once something is inside them, it’s infinitely reproducible โ€” a person can only be in one place, working for one employer, at a time. We’ve had a few hundred years to build tools for the human version of this problem: contracts, culture, non-competes. The model equivalent is still being invented in real time, which is exactly why Nadella felt the need to name it.

Apple’s recent legal action against former employees who joined OpenAI is this pattern in its sharpest form. Whatever the specifics, the shape is familiar: people who spent years inside one of the most sophisticated organizations in the world, carrying out knowledge that never appeared on any balance sheet and was hard to contain. No one fully anticipates what a mind absorbs simply by being in the room long enough.

That’s the real difference between the silicon case and the human one. You can try to take action to wall off knowledge flowing to a model. You cannot wall off what someone has learned to notice.

Categories
AI Apple Google

The Library You Already Own

Sharon Park in the morning is not a dramatic place. There’s a duck pond, a stand of oaks that go gold too briefly in November, and a loop I’ve walked enough times that my legs know it better than my eyes do. It is, in other words, exactly the kind of place where a person starts talking to himself. Not out loud. In the productive, low-grade way โ€” turning a sentence over, arguing with an idea from the day before, checking a thought against something you believe about yourself.

I think in five years I’ll be doing that walk with something else along. Not a search engine. Not another chatbot trained to know a little about everything and a lot about nothing in particular. Something closer to a second set of eyes on my own life โ€” a reasoning engine, lean and mostly private, that has actually read the things I’ve written and doesn’t need me to explain who I am before it’s useful.

Here’s the distinction that matters, and it took me longer than it should have to see it clearly. The AI industry has spent years in an arms race over how much of the world a model can hold โ€” more facts, more languages, more of the internet compressed into weights. That race will keep going, and somebody else can have it. What I want is smaller and stranger: a model that knows comparatively little about the world and quite a lot about me. My core values document. The portfolio spreadsheets. Fifteen years of blog posts. The half-finished notes for the I-280 project, sitting in a folder, waiting for someone โ€” or something โ€” to ask the right question about them.

I spent a career in payments infrastructure, which means I spent a career thinking about a very specific kind of trust: the kind where a stranger’s system has to make a judgment call, in milliseconds, about whether to say yes. Fraud models don’t work because they know everything about commerce. They work because they know an enormous amount about one account, one pattern, one person’s ordinary Tuesday โ€” enough to notice when Tuesday stops being ordinary. That’s the architecture I keep picturing, aimed inward instead of outward. Not a system trying to know the world. A system trying to know me, well enough to notice when I’m drifting from what I said I cared about.

I can already feel the shape of the mornings this would change. Right now, when I sit down to look at RMD requirements against the tax picture, I’m doing the translation myself โ€” pulling numbers into a story I can actually feel the weight of. A reasoning engine grounded in my real holdings wouldn’t just run the scenario. It would know that I don’t want the scenario dressed up as a spreadsheet; I want it dressed up as a conversation, unhurried, the kind you’d have over lunch with someone who already knows the whole situation. And on the mornings when I sit down to write, instead of staring at a blinking cursor and a blank page that has no idea I exist, I’d be handing a draft to something that has actually read my last two hundred posts and knows the difference between the sentence I’d write and the sentence I’d cut.

None of this is especially exotic technology. Apple and Google are already building toward it โ€” Neural Engines fast enough to do real reasoning on-device, retrieval systems that can reach into your own files instead of the entire internet, fine-tuning that’s getting cheap enough to personalize rather than merely customize. The more interesting story here isn’t privacy, though privacy is real. It’s architectural: what happens when the expensive, impressive part of the system โ€” the part that knows everything โ€” becomes optional, and the cheap, personal part โ€” the part that knows you โ€” becomes the whole point.

What I don’t yet know is what this will cost me. A tool that reasons this well about my own life is also a tool I could lean on instead of doing the leaning myself, and there’s a version of this future where the walk around Sharon Park stops being mine and starts being a conversation with something that finishes my sentences a little too well. I’d want some way of knowing, plainly, what it’s drawing from and what it’s guessing at โ€” less a nutrition label than a kind of honesty I could check against, the way you’d check a fraud model’s confidence score before you trusted it with a yes.

But most mornings, I think I’d take the trade. Not because I want to think less. Because for thirty years I’ve been collecting the raw material โ€” the notebooks, the portfolios, the half-built essays โ€” and it would be something, finally, to walk beside a mind that had actually done the reading.

Categories
AI

The Encyclopedia and the Reasoner

I was standing in the cereal aisle a few weeks ago, doing the thing I always do โ€” flipping the box over, scanning the fine print, comparing fiber grams like it mattered more than it probably does โ€” when I thought about the model I’d been testing that morning. Sharp. Fast. Occasionally, confidently, wrong about something I could have looked up in ten seconds.

There was no label for that. No panel telling me what was inside, what it was good at, what it might get wrong, what it cost to run. Just a chat window and a kind of blind trust.

That’s the itch behind this post. What would it look like if AI models came with something like a Nutrition Facts label โ€” the kind the FDA forced onto every box in your pantry back in 1994? Not as a gimmick, but as a real answer to a real problem: we are feeding these things into our decisions, our writing, our portfolios, our kids’ homework, largely on faith.

The IQ Number That Isn’t Quite an IQ Number

I keep running into a shorthand in investing circles โ€” Jordi Visser and others talking about frontier models as “140 IQ” systems, reasoning at a level that outpaces most humans on the kinds of puzzles we associate with fluid intelligence. Pattern recognition. Logic chains. Novel deduction under pressure.

It’s a useful number. It’s also a bit of a trick.

Human IQ tests were built to measure something narrow and specific โ€” not wisdom, not knowledge, not judgment, but the raw machinery of reasoning. When we borrow that language for AI, we inherit the same narrowness, which is fine as long as we remember it. A model that aces abstract reasoning benchmarks isn’t necessarily the model that knows the correct dosage, the right case law, or what actually happened in 1932. Reasoning and knowledge are cousins, not twins.

Two Kinds of Smart

Here’s an old-fashioned way to think about the split: Britannica versus World Book.

Britannica was the encyclopedia my father would have trusted โ€” dense, expert-written, unapologetically deep, assuming you could keep up. World Book was the one actually sitting on the shelf in most houses I knew growing up, mine included: friendlier, broader, built for a general reader, a little shallower in exchange for being a little more useful on a Tuesday night with a homework assignment due.

Neither is wrong. They’re optimized for different things. And training data does the same kind of sorting. A model fed heavily on curated, scholarly, expert-vetted sources leans Britannica โ€” deep, careful, occasionally slow to update. A model trained on the sprawl of the open web leans World Book โ€” broad, current, occasionally sloppy, sometimes brilliant at the edges precisely because it’s seen everything.

Any honest label for a model needs a section on this. Call it “Knowledge Sourcing.” Not just how big the training set was, but what kind of encyclopedia it’s pretending to be.

Sketching the Label

If I could design the box myself, it might read something like this:

Serving Size: 1 query, ~500 tokens

Reasoning Score: 138 (fluid problem-solving, logic, abstraction) Knowledge Depth: Moderateโ€“High (cutoff: [date]; strongest in [domains]; weakest in [domains])
Ingredients: Curated scholarly corpora, licensed news archives, public web crawl, synthetic reasoning data, human feedback Allergens: Confident hallucination under ambiguous prompts; recency gaps beyond training cutoff; known weakness in [specific domain]
Cost per Serving: $X per million tokens; Y watt-hours per query Best Paired With: Retrieval tools, human review for high-stakes decisions

It’s a little tongue-in-cheek written out like that. But underneath the joke is something I actually want โ€” the same instinct that made me read cereal boxes as a kid. Not to be scared of what’s inside, just to know.

The Part That Actually Excites Me

Here’s where the scaling laws get interesting, and where I think the real opportunity sits.

World knowledge is expensive. It’s greedy for data and parameters โ€” you need to have practically read the internet to know the boiling point of tungsten, the plot of a minor Victorian novel, and the org chart of a mid-cap company all at once. Reasoning, it turns out, is a different kind of animal. It can be distilled, compressed, taught through synthetic problems and careful post-training, and squeezed into something far smaller than you’d expect.

Which means a genuinely thrilling possibility is already taking shape: sharp, high-reasoning models small enough to run on a phone or a laptop, entirely offline, because they’ve shed the encyclopedia and kept the mind. Pair one of those with a personal index โ€” your own notes, your own documents, a retrieval layer built around your actual life โ€” and you get something closer to a personal thinking partner than a general-purpose oracle. Private. Fast. Always available. Tuned to you rather than to everyone. Apple may be on to something with this kind of strategy?

I think about this constantly in my own workflow โ€” the daily scans, the little agents I’ve built to help sort signal from noise, the genealogy digging, the investment frameworks I keep refining. What I usually want isn’t more encyclopedia. It’s a clear-headed reasoner sitting next to my own carefully kept knowledge, not buried under someone else’s version of the whole internet.

Why the Label Matters More Than the Score

None of this works, though, without honesty about what’s inside the box. A 140 on a reasoning benchmark tells you almost nothing about whether a model will quietly misremember a fact it was never that confident about in the first place. And a model can be extraordinarily knowledgeable while being a mediocre reasoner โ€” plenty capable of reciting the right ingredients and still getting the recipe wrong.

The nutrition label movement in food didn’t eliminate junk food. It just made it possible to choose junk food on purpose, with your eyes open, instead of by accident. I’d like the same deal with AI. Not a demand that every model be a genius generalist, but a demand that I get to know what I’m actually consuming โ€” and choose the lean local thinker over the bloated encyclopedia when that’s what the moment calls for, or the other way around when it isn’t.

Curiosity got me into that cereal aisle habit decades ago, and it’s the same instinct pulling me toward this idea now โ€” not suspicion of the box, just a wish to read it clearly before I decide how much of it to trust.

What would you want on your label?

Categories
AI Semiconductors

The Margin of the Weather

A company that has sold memory chips for forty years โ€” memory, one of the most humiliatingly commoditized products in capitalism, a business that has bankrupted entire Korean and Japanese conglomerates teaching each other lessons about discipline โ€” is about to make more money in twelve months than in the previous four decades combined.

Samsung’s chip chief told a room of his own employees: this year’s profit will exceed everything the division has earned since the 1970s. Forty years of grinding, erased by one fiscal year. You’d think they’d invented something.

They hadn’t. Everyone building an AI data center needs memory. Nobody built enough factories. Samsung was one of three companies on earth able to supply the shortfall, and the price of a chip that costs what it always cost went up fifty percent. Samsung kept the difference. Not innovation. What happens to a farmer when the drought hits every field but his.

We don’t credit the lucky farmer with genius. We say: good year. And we don’t expect the good year to repeat. Rain comes back. The price falls. Scarcity is weather, not a personality trait.

There’s a real achievement in this story too, and it has nothing to do with the weather. A year ago Samsung failed to qualify its most advanced memory for Nvidia’s systems โ€” performance problems, a rival getting the business instead. The engineers went back and fixed it. That’s the actual skill in this company’s year: unglamorous, uncelebrated at the town hall, worth nothing next to the number that got the confetti. The competence arrived quietly, on a different chip, in a different meeting, and nobody’s putting that on a plaque.

The stock market didn’t put it on one either, but it seemed to know the difference. Best quarter in Samsung’s history โ€” profit nineteen times the year before โ€” and the shares fell seven percent. Not despite the earnings. The gain had already been priced in, the shares having run up a hundred and fifty percent on the expectation of exactly this number, so the number’s arrival became a ceiling instead of a floor. A market rewards discovery. It does not reward weather. Had investors believed Samsung built something durable โ€” the Nvidia qualification, the years of engineering behind it โ€” the stock would have ripped, the way See’s Candies or Apple gets rewarded quarter after quarter, because everyone agrees the thing generating the money isn’t going anywhere. Instead the market glanced at the record harvest and asked, politely, whether it would rain again next year.

Analysts insist the shortage holds through next year. Someone always insists that, right before it doesn’t. Fabs get built. Capacity catches the demand that summoned it, the way it always has, and the cycle ends the way memory cycles end โ€” too much supply chasing too little demand, margins reverting toward the number they were always going to revert toward. Nobody knows if this time is different. A company just posted the best year of its life, on a windfall it didn’t earn and a fix it did, and the market โ€” which has seen droughts end before โ€” hasn’t decided yet which one it’s watching.

Categories
AI Business

The Wage of Knowing

In 1973 the Los Angeles Public Library installed a telephone line that worked while the building was dark. Dial H-O-O-T-O-W-L on a rotary phone, nine at night until one in the morning, and a librarian would answer. Somebody wanted to know the boiling point of mercury, or who wrote a poem they half remembered, or how many wives Henry VIII actually had, and a person on the other end of a cord found out. This went on for years. Nobody thought of it as data collection. It was just a service, a courtesy, a woman at a desk with a card catalog in her head.

I worked, in another life, in the payments industry, back when a merchant who wanted to charge your card had to call in and ask permission. There were rooms for this. Banks of phones, a bulletin of stolen numbers updated by hand, a floor limit past which a supervisor had to be found. The people answering the phones were, more often than you would guess, college students. Twenty years old, minimum wage, deciding in real time whether a stranger’s card was good. Nobody trained them for six months first. They learned the bulletin, they learned to listen for something wrong in a voice, and they said yes or no.

I have been driven, recently, by a car with nobody driving it. I noticed the wheel turning on its own and I braced for the wrongness of it. Thirty seconds later I was not bracing. I was looking out the window. The data says I was right to relax: across two hundred and twenty million miles, the cars involved in this experiment cause a small fraction of the serious crashes a human would have caused over the same roads. I did not need the data. I needed thirty seconds.

None of these people knew what they were doing. That is the thing about the librarian and the college student and, for that matter, about me learning to trust a wheel that moves by itself. The librarian was not building a search engine. The clerk was not training a fraud model. He was making rent. Their competence was not evidence, to them. It was just Tuesday. It became evidence later, to someone else, in a room they never saw โ€” the accident logs, the chargeback data, the accumulated record of a million correct guesses that turned out to be exactly the material a system needed to learn the job and take it.

This is the part that is easy to get wrong. It is not that the human failed and the machine succeeded. It is that the human succeeding, over and over, in full view, was the demonstration that the job could be learned. You do not automate a task nobody can do. You automate the one being done well enough, often enough, for long enough that the pattern becomes visible. Doing the job right was never neutral. It was the case being built.

Which brings me to a woman I will call the lawyer, because there are thousands of her and none of them are exactly her. She has a laptop open at her kitchen table. She logs into a dashboard belonging to a company that pairs credentialed people with the AI labs that need them โ€” a doctor here, a banker there, a corporate attorney with fifteen years of contract law behind her. She reads a model’s draft of a merger agreement and marks where it reasons like a first-year associate instead of a partner. She rewrites a clause. She explains, in the margin, why the model’s version would get laughed out of a negotiation. She is paid well for this. More, some weeks, than she billed certain clients.

She knows exactly what she is doing. That is the difference between her and the other three. The librarian did not know she was leaving a trail. The clerk did not know his good judgment would become someone else’s weights. I did not know, thirty seconds into that ride, that I was participating in anything at all. The lawyer knows. She is being paid, by the hour, at a rate that respects her expertise, to make her expertise legible enough that it no longer requires her. The company she works for has a name for this. They call it the reinforcement learning economy, which is a tidy way of saying: teach it everything, and then it will not need to call you back.

She does the work anyway. The rate is good. The work is interesting, in the way that teaching is interesting โ€” you learn what you know by trying to say it clearly enough for someone else to use. Nobody is lying to her. The dashboard does not pretend to be anything other than what it is. She logs off at the end of the session the way anyone logs off after a long day of being excellent at something, tired in the specific way that comes from careful work, and she does not, from what I understand, spend the evening thinking about what she has just fed into the machine.

I keep coming back to the rotary dial. Somebody dialing H-O-O-T-O-W-L at midnight in 1973 could not have imagined the lawyer at her kitchen table. But the shape is the same, if you look at it long enough. A person answers a question well. The answering becomes a record. The record becomes a system. The system answers next time. Nobody in the room ever decided this was the plan. It just turned out, every time, to be the plan.

Categories
AI Photography

The Price of the Cold

Two men are standing close to a brick wall trying not to talk, because talking wastes what little warmth is left in a body that has been outside too long. One of them has a camera โ€” Jerry Schatzberg, a fashion photographer. His hands are jammed half into his coat pockets between shots. The other man has his collar up around his ears and a scarf wound twice, black and white, and he is not moving much, because moving costs heat, and heat is the one thing neither of them has enough of. Schatzberg raises the camera. His fingers, by this point, are not entirely his own. When he presses the shutter there is a tremor in it he did not order and cannot undo.

The picture comes out smeared at the edges. Bob Dylan’s face, in the frame, is dissolving slightly into the gray behind him, like a man photographed through a windshield in the rain. It is, by any studio standard, a bad photograph. Schatzberg knows it’s a bad photograph. He has made a career out of not taking bad photographs.

And it became the cover of Blonde on Blonde, which is the best rock album ever recorded, and in nearly sixty years nobody has managed to improve on it by reshooting it clean. The blur isn’t a decision. It’s a symptom โ€” of two men standing in the cold too long, of a photographer choosing, afterward, to keep the evidence of his own discomfort instead of erasing it.

There’s a difference between an accident and serendipity that I don’t think gets said out loud enough, and it matters more than it used to. An accident is the cold โ€” involuntary, uninvited, spent before you know if it was worth spending. Schatzberg didn’t choose to shiver. His hands moved because his body was doing what bodies do at a certain temperature, and the shutter caught what his hands actually did, not what he meant to do. Serendipity is what happens next: a verdict, rendered after the fact, that the wreckage of an intention was better than the intention itself. The accident is what makes the verdict possible. Without the cold, there’s nothing to render a verdict on.

I’ve been sitting with a large language model most days for the better part of a year now, watching it write, asking it to try again, watching it try again in a way that is never quite the same and never quite different enough to matter. Somewhere upstream of me there is a number called temperature, and I will never see it. Somebody else did, once, in a meeting, and decided that the word for controlled, pre-approved, refundable randomness should be temperature โ€” the same word for the thing that made Schatzberg’s hands shake, the same word for the actual physical stakes of standing outside too long in January without enough coat โ€” and then set it, and moved on, and nobody in that meeting laughed, because nobody in the room had ever been cold in a way that mattered to the work.

Picture the room instead. It is climate-controlled to sixty-eight degrees, humidity held flat, year-round, by a building management system nobody thinks about until it fails. Somewhere in it, the hardware is generating your next five versions of a photograph like the one on Blonde on Blonde. Nobody in that room is going to lose feeling in their fingers today. Nobody’s collar is up. I don’t know his name โ€” nobody outside the building does โ€” but somebody like him tuned the sampling distribution and went home at six. That’s the guy in the good suit. He built the weather. He never once stood in it.

The small model inherits conclusions. It never inherits the cold. Whatever accidents shaped the teacher model’s own training โ€” whatever costly friction produced the insight in the first place โ€” the student model gets none of that weather. It gets the photograph, cropped and sharpened, with the blur removed because somebody along the way decided the blur was noise instead of signal โ€” the way Schatzberg, a lesser photographer, might have reshot Dylan clean and thrown the bad one away. It is heir to a serendipity it never earned, because it was never present for the accident that made the serendipity possible. It is, in the most literal sense the industry means by the word, cheap.

I keep coming back to the fact that nobody at the API layer is shivering. That’s not a complaint, exactly. It’s just an observation about where the cost went. Somewhere in the training data, some human being was cold, or scared, or holding a fish that was starting to smell, or standing on a stepladder with ten minutes before the traffic came back, and that person paid a real price for a result they couldn’t yet know was good. The model downstream of all that gets the result without the price.

Two rooms, then. In one of them it is January in New York and a man’s fingers have stopped entirely obeying him. In the other it is sixty-eight degrees, always, on a Tuesday and on a Sunday and at three in the morning, and the machines are making you nine more versions of that same blur. Sixty-eight degrees. A number, upstream, that you will never see.

Categories
AI

The Taste Beneath the Summary

The real work of staying informed has never been volume. It has been the quiet, repeated acts of judgment: does this matter, to whom, why now, what is the signal beneath the noise.

A recent piece from Bridgewater’s AIA Labs and Thinking Machines Lab, “Learning to Replicate Expert Judgment in Financial Tasks,” describes training models to do the triage investors actually doโ€”filtering news, research, central bank documents, internal notes, for relevance. Frontier models struggled with judgments that looked simple and weren’t. The fix wasn’t a bigger model. It was Qwen, fine-tuned on labeled examples from practitioners, and it beat the frontier leaders while costing a fraction to run.

The bottleneck was never model size. It was taste. And taste, it turns out, can be taught to something small and cheap, if you’re precise enough about what you’re teaching itโ€”a market’s worth of Mercors is already proving the same thing at scale.

The researchers were clear that expert judgment doesn’t reduce to rules or prompts. It took high-quality, domain-specific labels from people doing the actual work. The most powerful systems will be built in partnership with practitioners who can say, and keep saying, what “good” looks like in their own context.

Which raises the question I haven’t answered yet: what would I actually put in the labels, if someone asked me to teach my own taste to a cheap model.