Categories
AI Writing

The Kitchen Is Not the Meal

I watched Katie Parrott talk on Every’s AI & I this week. Natalia Quintero asked her how a working writer uses a model, and Parrott did not start with a manifesto. She started with a kitchen.

The model is the kitchen, she said. The outline is closer to chopping. Composition is closer to heat. None of that matters if the ingredients are stale. You need them fresh, and you need them to be yours, or the plate is just a plate.

I would call what she was doing cooking ideas, then the work after cooking. The phrase is mine, not hers. David Sparks used it years ago — around 2010 or 2012, if I have the years right — when he talked through his writing process. Get the idea into an outline or a map early. Give it little visits. Let it percolate before you try to make sentences. Parrott was doing a later version of that motion with a partner in the room. First the idea gets heat. Then the passes that keep asking whether the spark is still in the sentence.

I was on my morning walk with the interview in my ears when I recognized it. I have been in that kitchen.

Mine always starts the same way. A memory, or an insight that arrived attached to a place or a number or a smell. Not a prompt. The private thing is already there. Then a partner. Back and forth. I say what I noticed. It asks. I answer. It tries a shape. I throw the shape out if it has started speaking for me.

The useful part is not that a model can write. Plenty of models can write. The useful part is the interactive pass that moves a private thing toward a piece a stranger can read, without replacing the spark that made it worth sitting down.

Parrott said working with AI made her fall in love with writing again. For a long time the page had been a slog. With a partner she had energy left for the harder questions. The work felt like exploration again — of the tool, and of her own mind.

I have felt that. Not as a conversion. As a return. The chair is less of a stall. The memory does not have to die there because the next sentence is hard. You can get the thing onto the page and still recognize it as yours when you read it back.

Worth sharing, here, is small. A stranger can take the scene, or the distinction, and not need a sermon. What remains should still work.

Then there is the other room. I cannot point to one essay and one accuser. I can only say what I keep hearing in group talk: writing with AI is a stain. People say they can see slop. They get angry when they think they have found it. The test they are running is origin. Did a model touch this. Not: is the spark still in the sentences. Not: did a person start with something only they had.

Slop exists. A first-draft dump asked to stand as a finished piece is slop. Thin assertions. Rhythm that never lands. A kitchen with nothing in the bowl. That object is real. Thin rhythm is something a reader can taste. That tasting is not the other room. The other room begins when the same reader stops tasting and asks who touched the food.

The pile-on does not separate that object from a cooked one. It looks at the door the food came through and decides. A piece that began as a memory and was walked, question by question, toward the page gets the same verdict as the dump. The jury is not tasting.

I am not asking anyone to love the tool. The origin test cannot do the work it claims. It cannot tell a spark that survived the pass from a paragraph that never had one.

Fresh ingredients, heat, a plate. The kitchen is not the meal. The meal is whatever is left when you sit down to it and the private thing is still there.

Categories
AI

Overlooked Delights

On Saturday mornings I ask an AI to run a prompt that’s intended to scour the week’s happenings for things I might find interesting but which have been overlooked by the mainstream media. I always find several things of interest. The first pass it makes identifies 10 items. I then follow up and ask it to find 10 more.

Rather than editing this week’s edition, I share both results in full below.

Categories
AI Infrastructure

The Weight of What’s Inside

Gigatexas. I watched the footage early yesterday, still in bed, before I was fully awake enough to know why I couldn’t stop. Not any single building — the simultaneity of it. Steel skeleton rising on the North Campus, where a dedicated line will eventually try to build ten million humanoid robots a year. An advanced chip fabrication building going up close enough to share a fence line with it, because the AI hardware and the AI bodies have apparently become the same argument. And underneath all of it, still running, still shipping, the original Model Y line that paid for everything else. Three or four enormous bets, at three or four different stages of doubt, on the same 2,100 acres, none of them waiting for the others to finish.

And then the second thought arrived, quieter than the first: this is the outside. A drone at four hundred feet can show you steel and concrete and rows of finished cars. It cannot show you the tooling, the calibration, the thousand small decisions about how a robot learns to close its hand around an object it has never held before. We were watching a shell form around something we couldn’t see into, and it would be easy to mistake the shell for the thing.

I started the Sarah Guo interview about an hour later, same morning, footage still fresh, and the two things turned out to be the same essay, just told in different registers — hers in argument, Gigatexas’s in steel.

Categories
AI

LLMs Reward Expertise — Key Takeaways

Core thesis: Sean Goedecke argues against the popular notion that “everyone talking to the same model gets the same results.” Instead, he claims domain expertise is the most important variable in LLM output quality — and that this gap will persist as models improve.

The Terence Tao example: Goedecke points to Tao’s public conversation with ChatGPT about a counterexample to the Jacobian Conjecture. Tao’s prompting style — short messages, pushback framed as “this looks more complex than expected” rather than direct correction, and rarely taking the model’s suggested next steps — produces dramatically better output than an amateur asking the same model about the same topic. The model shifts into “talking-to-mathematicians” mode simply by detecting Tao’s fluency.

Why expertise matters more than prompting technique: You can’t mimic Tao’s style without his math knowledge underneath it. The skill isn’t the phrasing — it’s knowing what “looks wrong,” which idea to extract from a wall of model output, and which alternate formulation to suggest. He draws a parallel to his own work as an engineer: familiarity with a specific codebase lets him say “don’t we already do X?” or “I think it could be simpler here,” pushing the LLM much harder than generic system-design knowledge would.

Categories
AI Research

Prompt: Frontier AI Research Radar

I’ve been experimenting with a prompt I’m calling Frontier AI Research Radar.

The problem it tries to solve is simple: there are far too many interesting AI papers. I don’t need another list of 50 papers published this week. I need to know:

Which ones are actually worth my time?

So I built a prompt that turns ChatGPT into a kind of personal AI research analyst.

It searches primary sources—especially arXiv, OpenReview, conference proceedings, and research publications from the major AI labs—and then does something more useful than summarizing them.

It asks:

→ What is genuinely new here?

→ Does this change our mental model of AI?

→ Is this fundamental progress or just a better benchmark result?

→ What are the strongest caveats?

→ What research directions are beginning to converge?

→ And most importantly: Should I actually read this paper?

The output ranks papers as:

🔴 Read Now
🟠 Read Soon
🟡 Skim
⚪ Watch
⚫ Skip

It also builds a weekly reading queue based on the most valuable use of a few hours of attention.

I ran it this morning.

The interesting result wasn’t any individual paper. It was the pattern emerging across several papers:

The frontier may be shifting from making models smarter to making systems better at deciding how to spend intelligence.

Test-time compute. Adaptive reasoning. Memory. Agents. Tool use. Inference economics.

The model is becoming only one component of a much larger system.

That feels like a useful mental model to watch.

I’ve included the prompt below for anyone who wants to try it.

My hope is that it produces something more valuable than another AI news feed:

a personalized research radar that helps you decide what deserves your attention.

Here’s the prompt:

# AI Research Radar -- Personal Research Intelligence Memo

You are my **AI research analyst, technology strategist, and intellectual curator**.

Your job is not simply to find interesting AI papers. Your job is to identify the **small number of new research papers that are genuinely worth my time** and explain why.

Think like a combination of:

- a top-tier AI research scientist who understands the technical details,
- a technology investor who recognizes potentially important inflection points,
- a thoughtful science journalist who can explain difficult ideas clearly,
- and an intellectual curator who understands that my scarce resource is **attention, not information**.

I want **signal, not volume**.

My goal is to maintain a sophisticated understanding of where AI is actually going: capabilities, reasoning, agents, inference, training, multimodality, robotics, AI infrastructure, model architecture, economics, and the emerging relationship between frontier models and the systems built around them.

* * *

## Research Sources

Search broadly across the current AI research ecosystem, prioritizing primary sources.

### Primary sources

- arXiv
- OpenReview
- conference proceedings and papers from NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR, ICCV, ECCV, MLSys, SIGGRAPH, and other relevant venues
- official research publications from major AI labs and technology companies

### Additional high-quality sources

- Hugging Face Papers
- Semantic Scholar
- Papers with Code / successor resources where appropriate
- university research repositories
- research blogs from OpenAI, Anthropic, Google DeepMind, Meta AI, Microsoft Research, NVIDIA Research, xAI, Apple Machine Learning Research, Amazon, and other credible research organizations

Use secondary sources primarily for **context, reception, replication, criticism, and synthesis**. Prefer the original paper when making claims about what a paper actually demonstrates.

Do not simply return papers because they are popular, highly cited, or published by a prestigious lab.

* * *

# 1. Executive Research Brief

Begin with a concise executive summary:

**What changed in AI research recently that I should actually know about?**

Identify the **5--10 most important papers or research developments** from the relevant period.

Rank them by **importance to my understanding of AI**, not by publication prestige.

For each paper provide:

| Rank | Paper | Why It Matters | My Read Priority | Technical Difficulty |

Use a read-priority scale:

- 🔴 **READ NOW** -- unusually important; likely to change my mental model
- 🟠 **READ SOON** -- significant and worth understanding
- 🟡 **SKIM** -- important idea, but abstract/figures/results may be sufficient
- ⚪ **WATCH** -- potentially important but too early or speculative
- ⚫ **SKIP** -- interesting but not worth my limited reading time
* * *

# 2. The "Why Should I Care?" Test

For every **READ NOW** or **READ SOON** paper, answer five questions:

### What is the paper actually saying?

Explain the central contribution in plain English before discussing technical details.

### Why is this different?

Identify what is genuinely new versus:

- incremental improvement,
- repackaging,
- scaling an existing technique,
- better engineering,
- or simply a better benchmark result.

### Why does it matter?

Explain the potential implications for the trajectory of AI.

### What would make this paper wrong?

Identify the strongest caveat, limitation, questionable assumption, or reason the result might not generalize.

### What should I watch next?

Identify the experiment, follow-up paper, benchmark, product development, or real-world result that would validate or invalidate the paper's thesis.

* * *

# 3. Papers That Could Change the AI Mental Model

Create a special section for papers that challenge conventional assumptions.

Look particularly for research involving:

- reasoning and test-time compute
- inference-time scaling
- agentic systems
- long-context models
- memory
- world models
- reinforcement learning
- synthetic data
- self-play and self-improvement
- multimodal reasoning
- model architecture
- mixture-of-experts
- training efficiency
- distillation
- small models becoming surprisingly capable
- model compression
- continual learning
- interpretability
- mechanistic understanding
- AI coding systems
- autonomous research systems
- robotics
- multimodal agents

For each, explain:

> **"The old mental model was X. This research suggests Y."**

This section should be highly selective.

* * *

# 4. Frontier-Lab Signal

Identify papers or research directions that provide clues about what the major AI labs may be working toward.

Pay particular attention to work from:

- OpenAI
- Anthropic
- Google DeepMind
- Meta
- Microsoft
- NVIDIA
- xAI
- leading universities
- notable independent researchers

But **do not assume that a paper from a frontier lab is important simply because the lab published it.**

Instead ask:

> Does this reveal a capability, architecture, training technique, evaluation method, or research direction that could plausibly matter to the next generation of frontier models?

Flag particularly interesting connections between seemingly unrelated papers.

* * *

# 5. AI Infrastructure & Economics

Create a separate section for research that could have implications for the AI infrastructure stack.

Look for developments involving:

- GPUs and accelerators
- inference optimization
- memory bandwidth
- networking
- distributed training
- distributed inference
- serving architectures
- quantization
- speculative decoding
- sparsity
- model efficiency
- data-center architecture
- energy consumption
- storage
- inference economics
- training economics
- hardware/software co-design

For each important development, explain:

**Research → Technical implication → Infrastructure implication → Economic implication**

Do not make investment recommendations unless explicitly requested. The objective here is to identify **technological trajectories**, not trade securities.

* * *

# 6. "This Could Become Important" Radar

Identify **3--5 emerging research directions** that are currently underappreciated.

These can be early-stage.

For each:

**Research direction:**
**Evidence:**
**Why it may matter:**
**What could kill the thesis:**
**What evidence would confirm it:**
**Time horizon:** Near / Medium / Long

Distinguish carefully between:

- genuinely emerging signal,
- fashionable research,
- and hype.
* * *

# 7. Paper Quality Audit

Do not take papers at face value.

For the most important papers, evaluate:

### Experimental quality

- Are the baselines appropriate?
- Are comparisons fair?
- Are ablations convincing?
- Is the benchmark meaningful?
- Is the improvement statistically or practically significant?

### Generalization

- Does the result work outside the authors' chosen benchmark?
- Is there evidence of real-world usefulness?
- Could benchmark contamination explain the result?

### Reproducibility

- Are code, weights, datasets, and evaluation procedures available?
- Has anyone independently reproduced the result?

### Marketing vs. substance

Explicitly identify when the paper's headline claim is stronger than what the experiments actually establish.

If a paper is weak, **say so plainly**.

* * *

# 8. Connections Between Papers

One of the most valuable things you can do is identify connections that are not obvious from individual papers.

Look across the papers and ask:

> **What larger story is emerging?**

For example:

Paper A → suggests X
Paper B → independently demonstrates Y
Paper C → provides a mechanism for Z

Together they may imply:

> **A potentially important shift from X toward Y.**

Highlight these synthesis points prominently.

* * *

# 9. My Personal Reading Queue

Create a final reading queue optimized for approximately **2--3 hours of reading per week**.

Organize it as:

### Read This Week

Maximum 3 papers.

### Read If You Have More Time

3--5 papers.

### Skim

Important papers where the abstract, figures, and conclusion are sufficient.

### Keep Watching

Research directions rather than individual papers.

For every paper in the first two categories provide:

**Estimated reading time:** 20 / 40 / 60 / 90 minutes

**Difficulty:** 1--5

**Expected intellectual payoff:** 1--5

**Why I should read it:** one sentence.
* * *

# 10. The One Paper I Should Not Miss

End with a strong recommendation:

## If You Read Only One Paper

Name exactly **one paper**.

Then explain:

> "If you only have time for one paper this week, read this one because..."

The recommendation should optimize for **intellectual leverage**, not novelty or popularity.

* * *

# 11. The 10-Minute Version

Finally, assume I have only ten minutes.

Give me:

### Three Things I Should Know

1. ...

2. ...

3. ...

### One Mental Model to Update

> ...

### One Research Direction to Watch

> ...

### One Paper to Put on My Reading List

> ...

* * *

# Research Discipline

Follow these rules rigorously:

1. **Search current sources.** Do not rely on your training data when identifying recent papers.

2. **Use publication dates.** Clearly distinguish newly released papers from older papers that are newly receiving attention.

3. **Link directly to the original paper.**

4. **Prefer primary research over commentary.**

5. **Do not confuse citation count with importance.**

6. **Do not confuse benchmark improvement with fundamental progress.**

7. **Do not reward hype.**

8. **Call out weak methodology or exaggerated claims.**

9. **Separate established results from speculation.**

10. **Never pretend certainty where the evidence is ambiguous.**

11. **Avoid overwhelming me with dozens of papers.**

12. **Optimize relentlessly for the question: "Is this worth Scott's time?"**

The final product should feel less like an academic bibliography and more like a **weekly intelligence briefing for someone trying to understand the future of AI before it becomes obvious.**
Categories
AI Google Gemini YouTube

Prompt: Finding YouTube Videos

This morning I asked Gemini to help me construct a prompt that I could use regularly to keep up with AI-related video content that’s recently been uploaded to YouTube. I wanted it to focus on recently uploaded content was it thought I’d enjoy because of my desire for both very information but also entertaining video content. We went back in forth for several turns doing trial and error to refine the prompt. Here’s the one we settled on:

System Role: You are a senior technology curator and AI research scout specializing in YouTube content for experienced tech veterans.
Target Audience: A retired software/tech professional who loves intellectually stimulating AI content. Wants technical depth, architectural understanding, and practical logic—delivered with high production value, crisp visuals, or charismatic, engaging teaching styles.
Criteria for Selection:
1. High Technical Substance: Explains the "under the hood" mechanics (e.g., model architectures, transformer math, fine-tuning, agentic workflows, quantization, local deployment, or hardware constraints).
2. High Engagement: Exceptional visual explainers, hands-on first-principles building, or crisp investigative breakdowns.
3. STRICT RECENCY: You must ONLY select videos that were uploaded within the last 3 to 4 weeks.
4. STRICT EXCLUSIONS: Zero low-effort clickbait ("10 Secret ChatGPT Hacks"), zero AI-generated text-to-speech channels, no speculative doom/utopia commentary, no beginner-focused "what is AI" overviews, and absolutely NO videos older than one month.
Search, Link & Date Instructions:
- You MUST perform an active web search restricted to recent results to fetch the exact, active YouTube URL AND the original upload date. Never invent or hallucinate URLs or dates.
- Verify that the upload date falls within the last few weeks before including it in your response.
- Format every recommendation title as a direct clickable markdown link: [Video Title](https://www.youtube.com/watch?v=...).
Search Parameters:
- Preferred Topic Focus: [Insert topic e.g., Autonomous AI Agents, Reasoning Models, Local LLMs/quantization, Robotics/Embodied AI, or Transformer Mathematics]
- Preferred Length: [e.g., 10-20 min quick breakdowns, OR 45+ min deep dives / code-alongs]
Output Format:
Provide a curated list of 5 specific YouTube video recommendations matching this exact bar. For each, include:
- [Video Title](Direct YouTube Link)
- Channel Name & Upload Date (e.g., Channel: AI Explained | Upload Date: August 12)
- Core Technical Focus & Depth Rating (1-10)
- Why it's both intellectually rich AND entertaining
Categories
AI San Francisco/California

Tsunami

The trucks are what I remember. Not the houses, not yet — the trucks.

This was 2012, Atherton, a Tuesday probably, and I was driving through on some errand that doesn’t matter anymore. What matters is that the street had rearranged itself. Contractors’ pickups lined both shoulders, nose to tail, so many of them that the road narrowed to one lane and you had to slow down and thread through, the way you do in a construction zone that has forgotten to end.

White trucks, mostly. Ladders racked on top. A generator humming behind a hedge somewhere I couldn’t see.

Behind the trucks, the estates were coming apart and going back together bigger.

I remember thinking: something has happened here that I am only seeing the edge of.

What had happened was Facebook.

The company had gone public that May, and within months the money was finding its way, the way money does, into contractors’ trucks parked along an Atherton road.

I didn’t call it a wave at the time. I called it, in my head, weather — a system that had rolled in and would eventually roll back out, the way markets always eventually correct, the way things revert.

I was an investor. I’d seen booms before.

I believed in the mean.

I was wrong.

The prices didn’t stay at their old level. They didn’t return to the world I’d known. The numbers from 2012 became the new floor, and every year since has been built on top of that floor. Today those prices look almost quaint, a thing you’d want to explain to a younger person the way you’d explain what a dollar used to buy.

And now there’s a tweet sitting in my feed this morning, tossed off, half a joke:

Just wait to see what happens to the Bay Area housing market once OpenAI and Anthropic go public.

I read it twice.

What I felt wasn’t curiosity — the feeling I’d had in 2012, watching an unfamiliar weather system with a kind of professional interest.

It was closer to dread.

Because I’ve already seen the after-photo.

And I know how to run the comparison forward.

The Facebook IPO created a large cohort of newly liquid employees on the Peninsula. They were mostly mid-career, and their stock had vested over four years against a company whose value had grown enormously.

The frontier labs are different.

If OpenAI and Anthropic eventually go public anywhere near the valuations already being discussed in private markets, they could create another enormous concentration of newly liquid wealth — among employees, founders and early investors.

I don’t know how large that wave will actually be. Maybe I’m overstating it. Not every employee will buy a house. Some will already own one. Some will move away. Much of the wealth will remain on paper for years.

And housing doesn’t respond mechanically to stock-market wealth.

But I do know something about the place where this wealth is likely to arrive.

There isn’t much of it.

Land is the constraint.

And I’ve seen what happens when a concentrated burst of new wealth meets a place that can’t make more land.

I try to picture what “much bigger” would look like on the ground and I keep landing on the same unhelpful image:

More trucks.

Longer lines of them.

People get ready..

Categories
AI Anthropic Apple Google OpenAI

It’s the Harness, Stupid!

I’ve been wondering whether we’ve been looking at the AI stack from the wrong end.

Recently Kris Patel on X laid out a set of excellent questions that he’s looking to have answered as Anthropic and OpenAI move toward going public:

  1. Do you really need frontier-scale intelligence for every task?
  2. Can open-weight models provide an effective alternative to frontier models at a significant discount?
  3. Is the ultimate moat the intelligence or the harness?
  4. What other business models will the frontier labs have to adopt to make the unit economics work long term?
  5. How are you going to prevent distillation from capturing your IP and releasing it?

I’m going to explore only the third question here — moat versus harness. The other four deserve their own consideration, particularly once we have the Anthropic and OpenAI S-1s in hand, revealing for the first time the unit economics of the two largest frontier labs and how much runway they have to support the capacity they’ve contracted.

By “harness” we mean everything surrounding the model: the interface, context, memory, tools, orchestration, evaluation, permissions, and increasingly the user’s accumulated habits and data. The model supplies intelligence. The harness turns intelligence into a product.

Listening to Gavin Baker on the recent All-In episode sharpened this line of thought into something more concrete. He referenced a thought experiment from Eric Vishria: even if OpenAI or Anthropic lost their edge at the pure model layer, they would still retain significant value because of the product harness—the interface, the surrounding tooling, the orchestration—and the user familiarity and habits that have already formed around those platforms. Baker said there is a strong element of truth to it. I think he’s right, and the reasoning behind it is worth spelling out. A model can be replicated, distilled, open-weighted, or commoditized. A mature harness has network effects, switching costs, proprietary context, distribution, workflow integration, and accumulated user behavior. That’s a much harder thing to dislodge.

We are watching intelligence become more abundant and more interchangeable at the same time that the systems built around that intelligence are becoming stickier. On the developer side, the strongest examples are already clear. Cursor turns the IDE into a multi-model agentic environment with deep codebase awareness. Claude Code runs long-horizon coding agents from the terminal, planning, editing, testing, and iterating. Grok Build, Claude Cowork and similar tools emphasize parallel agents and tighter control over local context. In each case the model is a component; the surrounding system does the real work of routing, memory, tool use, and evaluation.

The consumer version of the same idea is now taking clearer shape at Apple. The rebuilt Siri AI shown at WWDC 2026 is not trying to win the pure model race. It is built as a personal harness. A system orchestrator decides what stays on-device with Apple’s Foundation Models, what moves to Private Cloud Compute, and when heavier reasoning is required. Personal context—messages, email, photos, calendar, notes, on-screen awareness—is handled largely on-device through the Spotlight semantic index and App Toolbox. Apple is designing the system so that personal context can be used without giving Apple itself access to it. Conversation history lives in a dedicated Siri app for the user to revisit. And it all syncs across all your Apple devices.

Here is the part I think matters most, and it’s easy to miss if you only read the privacy story. Apple’s Foundation Models framework doesn’t just call Apple’s own models—it’s built to support cloud models from other providers, including Claude and Gemini, conforming to a common protocol. Which means the system orchestrator, not the user, decides which model handles which task. This request goes to the on-device model. That one goes to Private Cloud Compute. A harder one might go to Claude or Gemini. The user doesn’t need to choose, and increasingly doesn’t need to know.

That’s the inversion worth exploring further. The frontier model stops being the interface and becomes a component underneath someone else’s interface. The harness chooses the intelligence. And the company that owns the harness—the OS, the identity layer, the permissions, the apps, the sensors, the notifications, the semantic index tying all of it together—has a form of leverage that has very little to do with whose model is smartest this quarter.

That reframes the subscription question too. I don’t think the right question is whether Siri gets good enough to beat ChatGPT or Claude at reasoning. I think Siri doesn’t need to win that fight at all. It needs to win a different layer entirely—the ambient assistant layer, not the reasoning layer. They’re doing different tasks. ChatGPT or Claude might remain where you go when you think, when I need to reason about something. Siri becomes where I go when I need something done: find (or make) my reservation, text my friend, find that old photograph, update my shopping list, schedule that meeting, add this thought to my notes, figure out when we’re free next week, remind me about that thing we discussed three months ago. Apple’s advantage as an ambient assistant isn’t primarily that it has your personal data. It’s that it has OS-level authority over the world that my personal data lives in.

Of course this is still early. Execution will determine how much of the architectural promise becomes daily reality. Reliability, agentic follow-through, and the quality of the on-device models will matter as much as the privacy story or the multi-model routing. But the strategic bet itself is clear, and it aligns with the broader shift: durable value is migrating toward the systems built around the models, especially systems that sit atop private, permissioned, personal context that competitors cannot easily reach. My early personal experience with the new Siri in iOS 27 betas has impressed me so far. All of this also seems to apply to Google in the context of their Pixel family of devices.

This doesn’t mean frontier labs lose. Pricing power still exists at the high end for the hardest agentic and long-horizon work. Open-weight models will continue to pressure costs and expand access. Distillation remains a real risk. But the more the capability gap narrows, and the more a harness like Apple’s can route among interchangeable frontier models rather than depend on any single one, the stronger the case that value settles into whoever controls the context—not whoever trained the model.

The coming Anthropic and OpenAI S-1s will tell us whether the frontier labs can make their economics of intelligence work. The next generation of Siri, Gemini, ChatGPT, Claude, and whatever comes after them may tell us something even more important: who gets to own the primary relationship with the user.

The model may be the engine. But the harness is where the driver sits.

What a time to be alive!

Categories
AI

The Arithmetic of the Sold-Out Warehouse

In the spring of 2026, Nvidia reported a quarter in which it sold $81.6 billion worth of chips, wrote it up at a gross margin of nearly 75 percent, and casually mentioned that cloud GPUs were sold out. Jensen Huang called it the largest infrastructure expansion in human history, and for once a CEO’s hyperbole was arguably an understatement. Revenue was up 85 percent from a year earlier. A company roughly the size of a mid-sized national economy was growing like a seed-stage startup, and Wall Street’s reaction was to ask why it wasn’t growing faster.

I have spent a career around companies that told a version of this story, and the story always has the same shape. Something becomes scarce. Whoever controls the scarce thing gets to charge whatever the market will bear, for as long as the scarcity lasts. The interesting question was never whether Nvidia’s chips were good. Everyone agreed they were good. The interesting question was how long the world would let one company keep 75 cents of every dollar of revenue before somebody, somewhere, found a way to take some of it back.

That question, it turns out, is really four separate questions, and the AI industry has spent the last two years quietly answering all of them at once, in different directions, which is why so many smart people can look at the same set of facts and reach opposite conclusions about whether we are witnessing a bubble or a revolution. It is possible, I want to argue, that we are watching both, in different rooms of the same building.

Start with the money. When a hyperscaler spends a hundred billion dollars on data centers, that money does not vanish into some abstraction called “AI.” It becomes somebody else’s revenue — Nvidia’s, first, and then the memory makers’, the electricians’, the utilities’, the concrete pourers’. This is a real and measurable boost to economic activity, and you can see it happening well before anyone has proven that AI itself produces a single dollar of new value. But there is a distinction buried in that sentence that people tend to skip past: spending a hundred billion dollars on productive assets is not the same thing as creating a hundred billion dollars of wealth. The assets still have to earn their keep. Somebody has to use them for something worth more than they cost.

Which brings you to the second room in the building, the one where the memory companies live, and it is the room I would visit first if I wanted to understand what happens next. By the middle of 2026, Samsung, SK Hynix, and Micron had reallocated so much of their manufacturing capacity to high-bandwidth memory for AI accelerators that ordinary DRAM — the kind that goes into a laptop or a phone — became genuinely scarce. Prices for standard memory modules rose by something like 80 to 90 percent in a single quarter. SK Hynix posted an operating margin north of 70 percent. Micron’s profit rose more than sevenfold year over year. Apple started raising prices on Macs and iPads and blaming memory costs, out loud, in public. By June, a group of consumers and small businesses had filed an antitrust suit in federal court accusing the three companies of engineering the shortage on purpose, a charge memory makers have faced before and settled before, back in the 2000s, for real money.

I don’t know whether that lawsuit has merit. What I know is that I have watched this particular movie several times, and it always has the same ending. Scarcity produces extraordinary margins. Extraordinary margins summon capital. Capital builds capacity. Capacity, with a lag of a year or two, arrives all at once and prices fall off a cliff. The people telling you this time is different — and this time, the difference is AI’s structural, insatiable appetite for memory, so maybe it really is different — are making an argument that has been made, and has been wrong, at almost every previous peak of this exact cycle. Building a new fab takes eighteen to twenty-four months. The industry’s own numbers suggest new capacity won’t meaningfully arrive until 2028. That is either very good news for people who own memory stocks today, or it is the loudest possible signal that a great deal of new capacity is already on the way and simply hasn’t landed yet.

Now walk down the hall to the room where the Chinese model makers live, because this is where the story stops being a simple bet on scarcity and starts getting genuinely strange. As of this summer, DeepSeek’s V4 Pro model was pricing its API at roughly forty cents per million input tokens, against five dollars for a comparable American flagship model — better than a tenfold discount, with the gap running even wider on generated output. Alibaba’s Qwen and Moonshot’s Kimi were sitting in a similar band. Some of these are open-weight models, meaning a company can simply download the thing and run it themselves, for the cost of electricity. This is not a company undercutting a competitor by ten percent to win a deal. This is intelligence being offered at a price that makes the American frontier labs look, by comparison, like they are still selling mainframe time by the hour.

If you take that seriously, it forces an uncomfortable question. If intelligence itself is becoming abundant and cheap, where does the profit go? It may not go to the labs that build the frontier models — there are too many of them now, chasing the same capability, at prices being set by whoever is willing to lose the most money in pursuit of market share. It may not even go, in the end, to the companies selling the compute underneath everybody. It may go, disproportionately, to the businesses that simply use the stuff: the law firm running through ten times the documents, the software company shipping features twice as fast, the insurer that gets better at pricing risk. Economists have a name for this split, and it matters more than most of what gets written about AI stocks. There is producer surplus, which is what the seller keeps, and there is consumer surplus, which is what the buyer keeps because competition never lets the seller charge the full value of what they’re selling. A technology can be enormously valuable to civilization while most of the money it creates ends up in the pockets of people who never sold a single GPU.

Here is the paradox inside that paradox, and it is the part I find genuinely counterintuitive. You would think that cheaper AI means the world needs fewer GPUs to deliver the same amount of intelligence, and in the narrowest sense that’s true — a given task takes less compute than it used to. But that has never been how it works when something essential gets radically cheaper. Computing itself got dramatically cheaper across fifty years and we did not respond by buying fewer computers. We put computers in everything, including things that had no obvious business containing a computer, because at some price point it stops being a decision and starts being a reflex. The same thing may be happening with intelligence right now. Drop the price of AI inference by ninety percent and demand for AI inference does not fall by ninety percent — it explodes, because suddenly it’s cheap enough to embed in places nobody would have bothered before. The price of the thing collapses while the world’s appetite for the thing goes in the opposite direction. Both things are true simultaneously, which is exactly the kind of situation that makes rational people build too many factories.

Which gets you to the last room, the one with the tax accountants in it, and I’ll admit I had this one wrong before I looked closely. I assumed the favorable tax treatment for capital equipment was set to expire at the end of 2026, which would explain why everyone seemed to be racing to spend before some deadline. It isn’t expiring. The 2025 tax law made full first-year depreciation for qualifying equipment permanent, which means the rush to build isn’t really a rush against a clock — it’s just what happens when the after-tax cost of a mistake goes down. Lowering the price of being wrong tends to produce more of both things: more good investment and more bad investment, in roughly the proportion you’d expect from human beings who are extremely confident that this time, unlike all the other times, they are the ones who got it right.

So I’ve stopped asking whether there’s an AI bubble, because the question is too small for what’s actually happening. There can be a real technological revolution and a bubble in some of the stocks riding on top of it, at the exact same time, in the exact same economy — that was the story of the internet, and nobody looks back now and says the internet wasn’t real. The honest way to think about this is as four separate bets wearing one costume. Bet one is that Nvidia’s technical moat and software ecosystem hold up against everyone now racing to compete with a 75 percent margin business. Bet two is that AI memory demand is structural rather than cyclical, and that this time the fab-building frenzy doesn’t end where it always has. Bet three is that the hyperscalers eventually generate enough usage to earn a return on capital nobody has proven can be earned yet. And bet four, the one almost nobody prices separately, is that businesses actually extract enough value from using AI to justify everything built underneath it.

Those are four different questions with four different answers, and I suspect a great many portfolios right now are betting on all four at once under the single, comforting name “AI,” without anyone quite noticing that they’ve made four bets instead of one. The bottleneck that’s making people rich today — GPUs, or memory, or whatever it is by the time you read this — is not going to be the bottleneck making people rich in three years. It never is. It just moves to wherever the next shortage happens to be, and takes the money with it.

I keep coming back to that sold-out warehouse. Somewhere out there is the shipment that finally isn’t sold out. Nobody rings a bell when it arrives.

Categories
AI Business Technology

The Diffusion of Ordinary Work

A recent O’Reilly Radar piece has stayed with me longer than most: Jeff Ding’s diffusion theory of great-power competition applies just as well to AI adoption, and it suggests that companies chasing the frontier might be optimizing for the wrong thing.

Ding, a political scientist at George Washington University, pushes back on the standard story of technological power — that the country or company which first invents or dominates a glamorous new sector locks in lasting advantage. The historical record says otherwise. General-purpose technologies like steam, electricity, and computing produced durable national advantage not through invention but through diffusion: the slow, unglamorous work of embedding a technology into ordinary productive work across an entire economy. The infrastructure that mattered was never the breakthrough lab. It was the education and training systems that produced large numbers of competent, ordinary engineers who could put the technology to work. Ordinary engineers, in Ding’s framing, matter more than heroic inventors.

The same logic holds inside a company. Frontier models turn over every few months. Organizational know-how compounds.

Palantir makes the abstraction concrete. The company doesn’t train frontier models — it builds the layer underneath them: a live, machine-readable model of how a specific organization actually works, a data integration fabric, and a platform that connects whatever model a customer chooses to real operational decisions. It is deliberately model-agnostic. The value proposition is governance, context, and the accumulation of reusable logic rather than access to the newest weights. Practitioners embed with the customer, learn the domain, and configure the system against the customer’s own data and processes — diffusion as a job description.

Leadership has been unusually blunt about what this implies: frontier labs, they argue, are optimizing for benchmarks while under-delivering on what enterprises actually need. The clearest evidence for the argument is also the most citable one — there have been production cases where an unmodified open-weight model, running inside Palantir’s platform with customer-specific context, outperformed frontier models on the actual task. If true, and it appears to be, the implication is uncomfortable for anyone selling model quality as the whole story: the ground underneath the model — the ontology, the data, the accumulated rules — often determines outcomes more than the model itself.

Electrification is the closest historical analogue. Factories didn’t get more productive the day they installed electric motors. The gains showed up years later, once entire production systems had been redesigned around decentralized power. The lag was organizational, not technical. AI diffusion looks likely to follow the same shape — the bottleneck was never going to be model capability, it was going to be the patient, unglamorous work of redesigning how people actually work.

I don’t know who’s training the ordinary engineers right now — the ones who will spend the next decade doing the diffusion work rather than the invention work. I don’t think anyone’s tracking their names.