Categories
AI Anthropic Apple Google OpenAI

It’s the Harness, Stupid!

I’ve been wondering whether we’ve been looking at the AI stack from the wrong end.

Recently Kris Patel on X laid out a set of excellent questions that he’s looking to have answered as Anthropic and OpenAI move toward going public:

  1. Do you really need frontier-scale intelligence for every task?
  2. Can open-weight models provide an effective alternative to frontier models at a significant discount?
  3. Is the ultimate moat the intelligence or the harness?
  4. What other business models will the frontier labs have to adopt to make the unit economics work long term?
  5. How are you going to prevent distillation from capturing your IP and releasing it?

I’m going to explore only the third question here — moat versus harness. The other four deserve their own consideration, particularly once we have the Anthropic and OpenAI S-1s in hand, revealing for the first time the unit economics of the two largest frontier labs and how much runway they have to support the capacity they’ve contracted.

By “harness” we mean everything surrounding the model: the interface, context, memory, tools, orchestration, evaluation, permissions, and increasingly the user’s accumulated habits and data. The model supplies intelligence. The harness turns intelligence into a product.

Listening to Gavin Baker on the recent All-In episode sharpened this line of thought into something more concrete. He referenced a thought experiment from Eric Vishria: even if OpenAI or Anthropic lost their edge at the pure model layer, they would still retain significant value because of the product harness—the interface, the surrounding tooling, the orchestration—and the user familiarity and habits that have already formed around those platforms. Baker said there is a strong element of truth to it. I think he’s right, and the reasoning behind it is worth spelling out. A model can be replicated, distilled, open-weighted, or commoditized. A mature harness has network effects, switching costs, proprietary context, distribution, workflow integration, and accumulated user behavior. That’s a much harder thing to dislodge.

We are watching intelligence become more abundant and more interchangeable at the same time that the systems built around that intelligence are becoming stickier. On the developer side, the strongest examples are already clear. Cursor turns the IDE into a multi-model agentic environment with deep codebase awareness. Claude Code runs long-horizon coding agents from the terminal, planning, editing, testing, and iterating. Grok Build, Claude Cowork and similar tools emphasize parallel agents and tighter control over local context. In each case the model is a component; the surrounding system does the real work of routing, memory, tool use, and evaluation.

The consumer version of the same idea is now taking clearer shape at Apple. The rebuilt Siri AI shown at WWDC 2026 is not trying to win the pure model race. It is built as a personal harness. A system orchestrator decides what stays on-device with Apple’s Foundation Models, what moves to Private Cloud Compute, and when heavier reasoning is required. Personal context—messages, email, photos, calendar, notes, on-screen awareness—is handled largely on-device through the Spotlight semantic index and App Toolbox. Apple is designing the system so that personal context can be used without giving Apple itself access to it. Conversation history lives in a dedicated Siri app for the user to revisit. And it all syncs across all your Apple devices.

Here is the part I think matters most, and it’s easy to miss if you only read the privacy story. Apple’s Foundation Models framework doesn’t just call Apple’s own models—it’s built to support cloud models from other providers, including Claude and Gemini, conforming to a common protocol. Which means the system orchestrator, not the user, decides which model handles which task. This request goes to the on-device model. That one goes to Private Cloud Compute. A harder one might go to Claude or Gemini. The user doesn’t need to choose, and increasingly doesn’t need to know.

That’s the inversion worth exploring further. The frontier model stops being the interface and becomes a component underneath someone else’s interface. The harness chooses the intelligence. And the company that owns the harness—the OS, the identity layer, the permissions, the apps, the sensors, the notifications, the semantic index tying all of it together—has a form of leverage that has very little to do with whose model is smartest this quarter.

That reframes the subscription question too. I don’t think the right question is whether Siri gets good enough to beat ChatGPT or Claude at reasoning. I think Siri doesn’t need to win that fight at all. It needs to win a different layer entirely—the ambient assistant layer, not the reasoning layer. They’re doing different tasks. ChatGPT or Claude might remain where you go when you think, when I need to reason about something. Siri becomes where I go when I need something done: find (or make) my reservation, text my friend, find that old photograph, update my shopping list, schedule that meeting, add this thought to my notes, figure out when we’re free next week, remind me about that thing we discussed three months ago. Apple’s advantage as an ambient assistant isn’t primarily that it has your personal data. It’s that it has OS-level authority over the world that my personal data lives in.

Of course this is still early. Execution will determine how much of the architectural promise becomes daily reality. Reliability, agentic follow-through, and the quality of the on-device models will matter as much as the privacy story or the multi-model routing. But the strategic bet itself is clear, and it aligns with the broader shift: durable value is migrating toward the systems built around the models, especially systems that sit atop private, permissioned, personal context that competitors cannot easily reach. My early personal experience with the new Siri in iOS 27 betas has impressed me so far. All of this also seems to apply to Google in the context of their Pixel family of devices.

This doesn’t mean frontier labs lose. Pricing power still exists at the high end for the hardest agentic and long-horizon work. Open-weight models will continue to pressure costs and expand access. Distillation remains a real risk. But the more the capability gap narrows, and the more a harness like Apple’s can route among interchangeable frontier models rather than depend on any single one, the stronger the case that value settles into whoever controls the context—not whoever trained the model.

The coming Anthropic and OpenAI S-1s will tell us whether the frontier labs can make their economics of intelligence work. The next generation of Siri, Gemini, ChatGPT, Claude, and whatever comes after them may tell us something even more important: who gets to own the primary relationship with the user.

The model may be the engine. But the harness is where the driver sits.

What a time to be alive!

Categories
AI Apple Google

The Library You Already Own

Sharon Park in the morning is not a dramatic place. There’s a duck pond, a stand of oaks that go gold too briefly in November, and a loop I’ve walked enough times that my legs know it better than my eyes do. It is, in other words, exactly the kind of place where a person starts talking to himself. Not out loud. In the productive, low-grade way — turning a sentence over, arguing with an idea from the day before, checking a thought against something you believe about yourself.

I think in five years I’ll be doing that walk with something else along. Not a search engine. Not another chatbot trained to know a little about everything and a lot about nothing in particular. Something closer to a second set of eyes on my own life — a reasoning engine, lean and mostly private, that has actually read the things I’ve written and doesn’t need me to explain who I am before it’s useful.

Here’s the distinction that matters, and it took me longer than it should have to see it clearly. The AI industry has spent years in an arms race over how much of the world a model can hold — more facts, more languages, more of the internet compressed into weights. That race will keep going, and somebody else can have it. What I want is smaller and stranger: a model that knows comparatively little about the world and quite a lot about me. My core values document. The portfolio spreadsheets. Fifteen years of blog posts. The half-finished notes for the I-280 project, sitting in a folder, waiting for someone — or something — to ask the right question about them.

I spent a career in payments infrastructure, which means I spent a career thinking about a very specific kind of trust: the kind where a stranger’s system has to make a judgment call, in milliseconds, about whether to say yes. Fraud models don’t work because they know everything about commerce. They work because they know an enormous amount about one account, one pattern, one person’s ordinary Tuesday — enough to notice when Tuesday stops being ordinary. That’s the architecture I keep picturing, aimed inward instead of outward. Not a system trying to know the world. A system trying to know me, well enough to notice when I’m drifting from what I said I cared about.

I can already feel the shape of the mornings this would change. Right now, when I sit down to look at RMD requirements against the tax picture, I’m doing the translation myself — pulling numbers into a story I can actually feel the weight of. A reasoning engine grounded in my real holdings wouldn’t just run the scenario. It would know that I don’t want the scenario dressed up as a spreadsheet; I want it dressed up as a conversation, unhurried, the kind you’d have over lunch with someone who already knows the whole situation. And on the mornings when I sit down to write, instead of staring at a blinking cursor and a blank page that has no idea I exist, I’d be handing a draft to something that has actually read my last two hundred posts and knows the difference between the sentence I’d write and the sentence I’d cut.

None of this is especially exotic technology. Apple and Google are already building toward it — Neural Engines fast enough to do real reasoning on-device, retrieval systems that can reach into your own files instead of the entire internet, fine-tuning that’s getting cheap enough to personalize rather than merely customize. The more interesting story here isn’t privacy, though privacy is real. It’s architectural: what happens when the expensive, impressive part of the system — the part that knows everything — becomes optional, and the cheap, personal part — the part that knows you — becomes the whole point.

What I don’t yet know is what this will cost me. A tool that reasons this well about my own life is also a tool I could lean on instead of doing the leaning myself, and there’s a version of this future where the walk around Sharon Park stops being mine and starts being a conversation with something that finishes my sentences a little too well. I’d want some way of knowing, plainly, what it’s drawing from and what it’s guessing at — less a nutrition label than a kind of honesty I could check against, the way you’d check a fraud model’s confidence score before you trusted it with a yes.

But most mornings, I think I’d take the trade. Not because I want to think less. Because for thirty years I’ve been collecting the raw material — the notebooks, the portfolios, the half-built essays — and it would be something, finally, to walk beside a mind that had actually done the reading.

Categories
AI Podcasts

A Remarkable Conversation…

Highly recommend this conversation between Harry Stebbings and Clay Bavor. Among many topics, I especially enjoyed the discussion about not investing in frontier models, the important values, the particular importance of craftsmanship, intensity, and family. And the special conversation about parenting and kids near the end. Just a delightful conversation to be able to enjoy!

Key Highlights:

• Founding Sierra: Bavor explains why he and Taylor chose to start Sierra, focusing on the transformative potential of language model-based agents (1:37 – 5:53).
• The AI Tech Stack: Sierra focuses on building enterprise-grade agent architectures and fine-tuning models on top of open-weights models rather than pre-training foundation models from scratch, prioritizing capital efficiency (5:53 – 7:15).
• Unbounded Demand for Intelligence: Bavor argues that there is massive, unmet demand for “frontier-level” intelligence in fields like coding, science, and legal work (7:15 – 11:41).
• Internal AI Operations: He details the use of Pinecone, an internal AI agent Sierra developed to navigate company data, streamline engineering, and assist in recruitment (18:36 – 22:00).
• Enterprise Strategy: Sierra employs a “forward-deployed” engineering model, embedding staff within client companies to ensure rapid, effective integration of AI, leading to quick deployment timelines (30:12 – 33:22).
• Board Governance: To keep pace with the speed of AI development, Sierra operates on a six-week board meeting cadence, utilizing comprehensive memos instead of traditional slide decks (39:07 – 41:13).
• Corporate Culture: Bavor emphasizes values like craftsmanship, intensity, and family. He also highlights the importance of working in-person to foster apprenticeship, mentorship, and a cohesive team culture (43:02 – 55:41).

Categories
AI Apple Google

The Floor

I compared the frontier to a three-star chef making grilled cheese in “Context Rot” — the smartest models on earth spending most of their time on work beneath them, the way a chef trained at Le Bernardin might still melt cheese between two slices of bread on a Tuesday night and call it dinner. The comfort was the point: if the sharpest tool is saved for hard problems and something merely-very-good handles the rest, nobody’s losing anything. The floor was never the interesting part.

I’ve kept turning the joke over, and I think I had the wrong worry.

Watch what companies do with their AI spend, not what they say. Coinbase moved engineers off frontier models onto open weights and cut its AI spend nearly in half while usage kept climbing. Nvidia runs a closed model as orchestrator and routes the actual volume — the daily uncelebrated bulk of it — to open weights it controls. The frontier is becoming a dispatcher, deciding where the request goes and rarely doing the work itself. The instinct is to worry about whose open weights end up running that volume, and right now the most capable ones at scale are Chinese — GLM, Kimi — which makes it tempting to read this as a contest America is quietly losing: the floor of the AI economy built somewhere else, at a price export controls can’t touch. You cannot embargo a file already downloaded. You cannot price-match free.

But that framing has a hole. Google’s own Gemma family is open-weight and good enough to handle that daily volume without anyone reaching for GLM or Kimi. “Open weights are a Chinese story” only holds if you don’t count the open models the company running Android and half the internet’s search traffic has already shipped.

And once I saw that hole, a bigger one opened behind it. I’ve been trying Apple’s new Siri — arriving with iOS 27 this fall, genuinely surprisingly good in beta — and it made me realize open weights, of any nationality, were never going to cook most of the world’s dinners. Apple and Google are.

Consider what actually determines where the world’s routine inference runs. Not which model benchmarks best, not which weights are downloadable — what’s already installed. Apple ships to well over a billion active devices before routing a single query through Siri’s new architecture. Nobody has to be persuaded to try it, or hear about it on a podcast; it’s the thing that answers when you press the button you’ve pressed for a decade. Google owns the search bar and the Android default the same way. Between them, that’s most of the world’s phones — and phones are where most of the world’s questions get asked.

The open-weight framing assumes the floor is up for grabs, that whoever ships the best free model wins the daily grind by merit. But the floor was never a bazaar. It’s a set of defaults, owned by whoever already has the device in your hand, not whoever holds the most generous license. Apple didn’t need to win the model war to win this. Its heaviest reasoning tier is built with Google, running on Nvidia chips in Google’s cloud, under a deal reported at roughly a billion dollars a year — Apple doesn’t fully own the engine doing the thinking. It doesn’t need to. It owns the button.

That’s a quieter concentration than an export-controls fight, and a harder one to dislodge. An open model can be forked, distilled, undercut, or out-competed by the next release. A billion phones with an assistant built into the lock screen cannot be routed around. Whoever’s weights hum underneath barely matters, the way it barely matters to a diner which supplier delivered the flour. What matters is whose kitchen the meal came from, and whose name is on the door.

The grilled-cheese chef was never the risk. Two chefs are about to own nearly every kitchen on earth, and most of us will never notice — because a kitchen you’ve been eating out of for a decade doesn’t feel like something that was won. It just feels like home.

Owning the kitchen and getting paid for what’s cooked in it, though, turn out to be two different questions. That one’s for another post.

Categories
AI Anthropic Business Google

The Weight of the Bill

Jordi Visser has been making the case for months — in his weekly YouTube commentary and on his Substack — that we are living through an exponential transition that most people are measuring with the wrong instruments. I think he’s right. I found two data points this week that suggest why.

I was somewhere in the middle of an Invest Like the Best episode when Dylan Patel said it — almost as an aside, the kind of thing you drop to establish context before moving on to the point you actually came to make. His firm, SemiAnalysis, analyzes the semiconductor and AI industries for a living. And their usage of Claude, he noted, has been growing. The costs have been growing too.

Exponentially.

He moved on. I didn’t.

I think Patel’s API bill might be one of the more honest documents in the current AI moment — more honest than the analyst reports his firm produces, more honest than the earnings calls where every public company performs its AI fluency for shareholders.

Surveys bend. When you ask someone whether they’re using AI in their work, you’re asking them to self-report on a technology that has become a proxy for relevance, for not being left behind. The incentive to say yes is enormous. And even when the yes is genuine, it tells you nothing about depth — whether AI has become load-bearing in how someone actually works, or whether it’s an impressive thing they do occasionally.

Nobody pays exponentially growing API costs for show. Money is the honest witness.

What makes Patel’s situation quietly strange is the recursion in it. SemiAnalysis exists to help sophisticated investors and technologists understand this industry — and they cannot predict their own consumption curve. They are inside the exponential the same way everyone else is. They just happen to be watching their bill.

Then this morning, a different number arrived. Google announced it will invest up to $40 billion in Anthropic — $10 billion committed now, another $30 billion contingent on performance milestones. This follows a separate $5 billion from Amazon, part of a broader arrangement under which Anthropic is expected to spend up to $100 billion on compute over time.

The temptation with numbers like these is to treat them as spectacle. Forty billion dollars is so large it becomes almost aesthetic — a statement about ambition, about the kind of bets that define eras. You feel the weight of the zeros and move on.

But I keep coming back to Patel’s API bill.

Because Google’s $40 billion and SemiAnalysis’s compounding monthly costs are saying the same thing, expressed at scales so different they almost don’t seem related. One is a research firm noticing that their tool usage has quietly escaped prediction. The other is one of the most sophisticated capital allocators on earth making a bet that strains comprehension. But both are pointing at the same reality: that this technology, wherever it takes hold, does not plateau. It compounds.

We have been waiting, I think, for the moment when AI adoption becomes legibly real — some threshold event that separates the signal from the noise, the press release from the actual change. The surveys were supposed to mark that moment. The enterprise announcements. The benchmark numbers.

Patel’s aside suggests we’ve been waiting for the wrong thing. You don’t arrive at the exponential. You just eventually notice you’re already in it — in an aside on a podcast, before moving on to the point you actually came to make.

Categories
AI

The Geometry of Speed

We are surprised when witnessing something move faster than our intuition expects. We are inherently wired to understand slow, compounding growth. We expect the long, grinding years of the plateau—the quiet periods where nothing seems to happen before a sudden breakthrough.

I was looking at a chart Patrick Collison shared this morning, and it challenged that very intuition. It’s a simple, stark visualization: AI model intelligence relative to the formation date of the lab that built it.

If you trace the lines for Google and OpenAI on the right side of the graph, you see the history we’ve all lived through. Thousands of days—more than a decade of quiet, methodical, often unglamorous research—before their trend lines finally bend and shoot upward. It is a geometry of patience. It’s the visual representation of laying bricks, one by one, year by year, until you have a foundation sturdy enough to support the weight of a revolution.

And then, on the far left of the chart, there is a red line. MSL. The team behind Meta’s new Muse Spark model, released today.

The red line doesn’t curve. It doesn’t slope. It simply strikes straight up, like a lightning bolt in reverse.

In roughly 200 days since formation, this new effort achieved a level of capability that took the early pioneers thousands of days to reach. Collison noted how much he loves seeing things done quickly, and it’s hard not to share that specific, visceral thrill of seeing the boundaries pushed so aggressively.

I find myself thinking about the architecture of speed and what it means for the rest of us.

We spend so much of our lives absorbing the lesson that “good things take time.” We are taught that the crucible of meaningful work requires a long, slow simmer. And mostly, that remains true. The compound interest of human experience is real, and wisdom is rarely rushed.

Yet, every once in a while, a new paradigm emerges that doesn’t just accelerate the timeline—it collapses it entirely.

The pioneers cut the agonizingly slow path through the jungle, taking the brunt of the time, the friction, and the missteps. The ones who follow—like xAI, Anthropic, and now MSL—don’t have to clear the brush from scratch. They can look at the map, pave the road, and simply drive.

What does it mean for our own mental models when the timeline from “formation” to “frontier” shrinks from five thousand days to a few hundred?

It is a jarring reminder that the past pace of performance is not a law of physics.

I think about my own assumptions—how often I assume a project, a habit, or a societal shift will take a while, simply because similar things took a while in the past. We anchor our expectations to old geometry.

Meta’s release of Muse Spark is a technical feat, certainly. But the chart itself holds a broader, more human lesson. It’s a visual prompt to constantly re-evaluate our assumptions about how long the impossible is supposed to take.

The future doesn’t always arrive on a comfortable, predictable schedule. Sometimes, it just shows up unannounced, demanding we adjust our stride to keep up.

Categories
Music

Every Blog Needs a Theme Song!

Google has added a new music generation model called Lyria 3 to its Gemini 3 models.

I was playing around with it last night – having it generate happy birthday greetings for a friend whose birthday is coming up in a few days, another song for a longtime business partnership I was part of, and more. It’s kind of crazy! And a lot of fun.

When you use Lyria 3 as a tool in Gemini 3 you get back an image and an MP3 file that’s 30 seconds long (longer coming soon according to Google). Turns out the 30 second length is just about perfect for the “quick hit” from a snippet of music.

Google provides several genres you can choose from to start with – or you can just go with whatever you want to say in the prompt – here’s a rough template for doing that:

[Topic] + [Genre] + [Mood] + [Instruments] + [Vocals]

This morning I went for my morning walk and had a thought – how about generating a theme song for my blog. So when I got back home I opened up Gemini, selected the Music tool and entered:

Take a look at my blog and compose my theme song! blog: https://sjl.us

You can see with that prompt that I really didn’t provide it much direction – just a pointer to my blog so that it could try to generate something appropriate.

It took a few seconds for Lyria to read my blog and then use what it found to generate my blog’s theme song – and I like it!

You can play the theme song for yourself here:

Categories
AI

The Second Fire: From Finding to Forming

There is a specific kind of vertigo that comes with a paradigm shift. It’s the feeling of standing on the edge of a map that has just been unrolled to reveal twice as much territory as you thought existed. Lately, as I navigate the vast, generative landscape of AI, that old vertigo has returned. It’s a hauntingly familiar resonance, a structural echo of the late nineties and early 2000s when we first encountered the Google search bar.

Back then, the world was a series of closed doors. Information was siloed in physical libraries, expensive encyclopedias, or the unreliable oral histories of our social circles. Then came that clean, white interface with a single blinking cursor. Suddenly, the friction of “not knowing” began to evaporate. We weren’t just browsing the web; we were suddenly endowed with a collective memory. It felt like a superpower—the ability to summon any fact from the digital ether in milliseconds.

“Google is not just a search engine; it is a way of life. It is the way we find out who we are, where we are going, and what we are doing.”

Today, the sensation is different in texture but identical in weight. If Google gave us the power to find, AI is giving us the power to form.

The “Aha!” moment of 2026 isn’t about locating a PDF or a Wikipedia entry; it’s the realization that the distance between a thought and its realization has shrunk to almost nothing. When I prompt a model to synthesize a complex theory or visualize a dream, I feel that same electric jolt I felt twenty years ago when I realized I’d never have to wonder about a trivia fact ever again.

But there is a philosophical weight to this new “awesome.” With Google, the challenge was discernment—filtering the flood of information to find the truth. With AI, the challenge is intent. When the “how” becomes effortless, the “why” becomes the only thing that matters. We are moving from the era of the Librarian to the era of the Architect.

We are once again holding a new kind of fire. It’s warm, it’s brilliant, and just like the first time we saw that search bar, we know that the world we lived in yesterday is gone, replaced by a version where our reach finally matches our imagination.

Categories
AI Books Google NotebookLM San Francisco/California Writing

The 280 Project

Way back in 2016 when I was contemplating my retirement, I found myself pondering what projects might keep me engaged once my long-standing career in payments consulting came to an end. One compelling idea that emerged during this reflective period was the prospect of writing another book. This time, I envisioned the topic focusing on the intriguing story behind Interstate 280, often referred to as “the world’s most beautiful freeway.”

Our family’s migration from the Midwest to California took place in the early 1960s, a time when the interstate highway system in the San Francisco Bay Area was still a work in progress. At that point, I-280 had not yet been completed. As I approached the age of obtaining my driver’s license and gained the freedom that came with access to a car, I remember setting off on explorative drives down the peninsula. During those excursions, I gradually became aware of the ongoing construction and development involved in building this iconic road.

Eventually, after years of planning and labor, I-280 was completed in the early 1970s. At that time, I was working for IBM and was engaged in a project that took me down to an IBM lab facility located on Sand Hill Road—a place that has since vanished. Driving along I-280 during those initial years was an absolute delight, with the smooth asphalt feeling fresh and new under my tires. The experience of traversing a well-constructed highway surrounded by natural beauty was euphoric.

Sidenote: that IBM lab on Sand Hill Road was where Gene Amdahl was working on what turned out to be his last project working for IBM. That project was abruptly terminated one day and Amdahl left to found what became Amdahl Computer, developer of the first of the serious IBM mainframe “clone” threats.

In stark contrast to other freeways that meander through urban landscapes or feature monotonous views, 280’s route is distinguished by its breathtaking scenery. The rolling hills, lush vegetation, and stunning vistas create a picturesque drive that sparkles in comparison to its sibling highway, US 101, which navigates through the more densely populated areas closer to San Francisco Bay.

As I brainstormed the possibility of transforming my interest in I-280 into a full-fledged book project, I realized there must be an abundance of fascinating stories to uncover regarding the history of this highway—particularly pertaining to how the route was established and agreed upon. To delve deeper into this narrative, I invested considerable time gathering a wealth of documents. A few hours of dedicated Google searches yielded a treasure trove of information, which I organized into a folder for easy access. However, I soon found myself lacking a clear methodology for effectively utilizing these documents to craft an engaging narrative.

Recently, I have begun experimenting with Google’s NotebookLM, which appears to be tailored precisely to meet my needs. This innovative tool allows me to input numerous documents and then facilitates various inquiries about the collected material. I can explore whether there are any captivating and compelling stories waiting to be told. As I embark on this new journey of exploration, I am filled with a sense of excitement and renewed vigor for my little project. While it remains uncertain whether a full-fledged book will emerge from this endeavor, I am intrigued by the possibilities and look forward to seeing how this story unfolds. Perhaps this exploration will not only breathe life into my ideas but also provide a narrative worth sharing with others. We shall see!