Categories
AI Anthropic Apple Google OpenAI

It’s the Harness, Stupid!

I’ve been wondering whether we’ve been looking at the AI stack from the wrong end.

Recently Kris Patel on X laid out a set of excellent questions that he’s looking to have answered as Anthropic and OpenAI move toward going public:

  1. Do you really need frontier-scale intelligence for every task?
  2. Can open-weight models provide an effective alternative to frontier models at a significant discount?
  3. Is the ultimate moat the intelligence or the harness?
  4. What other business models will the frontier labs have to adopt to make the unit economics work long term?
  5. How are you going to prevent distillation from capturing your IP and releasing it?

I’m going to explore only the third question here — moat versus harness. The other four deserve their own consideration, particularly once we have the Anthropic and OpenAI S-1s in hand, revealing for the first time the unit economics of the two largest frontier labs and how much runway they have to support the capacity they’ve contracted.

By “harness” we mean everything surrounding the model: the interface, context, memory, tools, orchestration, evaluation, permissions, and increasingly the user’s accumulated habits and data. The model supplies intelligence. The harness turns intelligence into a product.

Listening to Gavin Baker on the recent All-In episode sharpened this line of thought into something more concrete. He referenced a thought experiment from Eric Vishria: even if OpenAI or Anthropic lost their edge at the pure model layer, they would still retain significant value because of the product harness—the interface, the surrounding tooling, the orchestration—and the user familiarity and habits that have already formed around those platforms. Baker said there is a strong element of truth to it. I think he’s right, and the reasoning behind it is worth spelling out. A model can be replicated, distilled, open-weighted, or commoditized. A mature harness has network effects, switching costs, proprietary context, distribution, workflow integration, and accumulated user behavior. That’s a much harder thing to dislodge.

We are watching intelligence become more abundant and more interchangeable at the same time that the systems built around that intelligence are becoming stickier. On the developer side, the strongest examples are already clear. Cursor turns the IDE into a multi-model agentic environment with deep codebase awareness. Claude Code runs long-horizon coding agents from the terminal, planning, editing, testing, and iterating. Grok Build, Claude Cowork and similar tools emphasize parallel agents and tighter control over local context. In each case the model is a component; the surrounding system does the real work of routing, memory, tool use, and evaluation.

The consumer version of the same idea is now taking clearer shape at Apple. The rebuilt Siri AI shown at WWDC 2026 is not trying to win the pure model race. It is built as a personal harness. A system orchestrator decides what stays on-device with Apple’s Foundation Models, what moves to Private Cloud Compute, and when heavier reasoning is required. Personal context—messages, email, photos, calendar, notes, on-screen awareness—is handled largely on-device through the Spotlight semantic index and App Toolbox. Apple is designing the system so that personal context can be used without giving Apple itself access to it. Conversation history lives in a dedicated Siri app for the user to revisit. And it all syncs across all your Apple devices.

Here is the part I think matters most, and it’s easy to miss if you only read the privacy story. Apple’s Foundation Models framework doesn’t just call Apple’s own models—it’s built to support cloud models from other providers, including Claude and Gemini, conforming to a common protocol. Which means the system orchestrator, not the user, decides which model handles which task. This request goes to the on-device model. That one goes to Private Cloud Compute. A harder one might go to Claude or Gemini. The user doesn’t need to choose, and increasingly doesn’t need to know.

That’s the inversion worth exploring further. The frontier model stops being the interface and becomes a component underneath someone else’s interface. The harness chooses the intelligence. And the company that owns the harness—the OS, the identity layer, the permissions, the apps, the sensors, the notifications, the semantic index tying all of it together—has a form of leverage that has very little to do with whose model is smartest this quarter.

That reframes the subscription question too. I don’t think the right question is whether Siri gets good enough to beat ChatGPT or Claude at reasoning. I think Siri doesn’t need to win that fight at all. It needs to win a different layer entirely—the ambient assistant layer, not the reasoning layer. They’re doing different tasks. ChatGPT or Claude might remain where you go when you think, when I need to reason about something. Siri becomes where I go when I need something done: find (or make) my reservation, text my friend, find that old photograph, update my shopping list, schedule that meeting, add this thought to my notes, figure out when we’re free next week, remind me about that thing we discussed three months ago. Apple’s advantage as an ambient assistant isn’t primarily that it has your personal data. It’s that it has OS-level authority over the world that my personal data lives in.

Of course this is still early. Execution will determine how much of the architectural promise becomes daily reality. Reliability, agentic follow-through, and the quality of the on-device models will matter as much as the privacy story or the multi-model routing. But the strategic bet itself is clear, and it aligns with the broader shift: durable value is migrating toward the systems built around the models, especially systems that sit atop private, permissioned, personal context that competitors cannot easily reach. My early personal experience with the new Siri in iOS 27 betas has impressed me so far. All of this also seems to apply to Google in the context of their Pixel family of devices.

This doesn’t mean frontier labs lose. Pricing power still exists at the high end for the hardest agentic and long-horizon work. Open-weight models will continue to pressure costs and expand access. Distillation remains a real risk. But the more the capability gap narrows, and the more a harness like Apple’s can route among interchangeable frontier models rather than depend on any single one, the stronger the case that value settles into whoever controls the context—not whoever trained the model.

The coming Anthropic and OpenAI S-1s will tell us whether the frontier labs can make their economics of intelligence work. The next generation of Siri, Gemini, ChatGPT, Claude, and whatever comes after them may tell us something even more important: who gets to own the primary relationship with the user.

The model may be the engine. But the harness is where the driver sits.

What a time to be alive!

Categories
AI Business Technology

The Diffusion of Ordinary Work

A recent O’Reilly Radar piece has stayed with me longer than most: Jeff Ding’s diffusion theory of great-power competition applies just as well to AI adoption, and it suggests that companies chasing the frontier might be optimizing for the wrong thing.

Ding, a political scientist at George Washington University, pushes back on the standard story of technological power — that the country or company which first invents or dominates a glamorous new sector locks in lasting advantage. The historical record says otherwise. General-purpose technologies like steam, electricity, and computing produced durable national advantage not through invention but through diffusion: the slow, unglamorous work of embedding a technology into ordinary productive work across an entire economy. The infrastructure that mattered was never the breakthrough lab. It was the education and training systems that produced large numbers of competent, ordinary engineers who could put the technology to work. Ordinary engineers, in Ding’s framing, matter more than heroic inventors.

The same logic holds inside a company. Frontier models turn over every few months. Organizational know-how compounds.

Palantir makes the abstraction concrete. The company doesn’t train frontier models — it builds the layer underneath them: a live, machine-readable model of how a specific organization actually works, a data integration fabric, and a platform that connects whatever model a customer chooses to real operational decisions. It is deliberately model-agnostic. The value proposition is governance, context, and the accumulation of reusable logic rather than access to the newest weights. Practitioners embed with the customer, learn the domain, and configure the system against the customer’s own data and processes — diffusion as a job description.

Leadership has been unusually blunt about what this implies: frontier labs, they argue, are optimizing for benchmarks while under-delivering on what enterprises actually need. The clearest evidence for the argument is also the most citable one — there have been production cases where an unmodified open-weight model, running inside Palantir’s platform with customer-specific context, outperformed frontier models on the actual task. If true, and it appears to be, the implication is uncomfortable for anyone selling model quality as the whole story: the ground underneath the model — the ontology, the data, the accumulated rules — often determines outcomes more than the model itself.

Electrification is the closest historical analogue. Factories didn’t get more productive the day they installed electric motors. The gains showed up years later, once entire production systems had been redesigned around decentralized power. The lag was organizational, not technical. AI diffusion looks likely to follow the same shape — the bottleneck was never going to be model capability, it was going to be the patient, unglamorous work of redesigning how people actually work.

I don’t know who’s training the ordinary engineers right now — the ones who will spend the next decade doing the diffusion work rather than the invention work. I don’t think anyone’s tracking their names.

Categories
AI

Claude as Walter Cronkite

Gavin Baker said something this week that stuck with me.

In his latest conversation with Patrick O’Shaughnessy, he described a quiet shift happening across public markets. Nearly everyone he knows in the equity business—retail and institutional—now feeds every piece of news straight into Claude. Sometimes Claude Code. Sometimes a Claude agent. The model is probabilistic, he noted, and he was speaking from what he sees in his own network rather than from a measured study. But his impression was that the variation in how it interprets the same information is surprisingly small. A huge chunk of the market ends up trading on a shared reading of events.

Baker reached for an old analogy: Claude has become Walter Cronkite for the stock market. The single trusted voice. Everyone just believes what it says.

He tied the observation to Michael Mauboussin’s work on how a breakdown in diversity of thought helps create the conditions for bubbles and crashes. When independent judgment collapses into a narrower set of interpretations, the system becomes more brittle. Moves get sharper. Errors get amplified.

I spent the back half of my career inside fraud detection systems at Visa, watching correlated failure up close. The lesson that never left me: the dangerous moment isn’t when a single model is wrong. Individual errors wash out. It’s when every model in the ecosystem is wrong in the same direction, because they were trained on the same data, tuned against the same benchmarks, built by people reading the same journals and hiring from the same three schools. A fraud ring doesn’t need to beat your model. It needs to find the blind spot every model in the industry shares. That’s not a tail risk. That’s the whole risk.

Which is what made me sit up a few weeks ago, watching a position reprice in a straight line and catching myself, mid-scroll, about to ask Claude what it thought was happening before I’d looked at a single primary source myself. The tool hadn’t done anything wrong. I had reached for the shared interpretive layer before reaching for my own judgment, out of habit, the way you reach for a light switch in a dark room you’ve walked through a thousand times.

Dan Geer wrote about this two decades earlier, from a different angle entirely. Geer and colleagues argued that Microsoft’s dominance had created a software monoculture: nearly identical systems sharing the same vulnerabilities. In biology, monocultures are efficient until a pathogen finds the common flaw. Then the failure is systemic rather than local. Diversity limits the blast radius. Geer’s point was never that the dominant platform was worse in isolation. It was that identicality itself becomes the risk multiplier.

Baker is describing a cognitive version of the same phenomenon.

The platform is no longer Windows. It is a frontier model that a large fraction of market participants now use as their primary interpretive layer. The shared vulnerability is not a buffer overflow. It is a common set of priors, training data, reasoning patterns, and prompt conventions. Slight probabilistic differences still exist. But the center of gravity of interpretation has tightened.

The result is correlated positioning. Feedback loops that reinforce themselves. A market that can reprice more violently than the underlying fundamentals alone would justify. In July we watched AI and semiconductor names drop 40–60 percent in a straight line while on-the-ground metrics—GPU rental prices rising, token growth accelerating, hyperscaler operating cash flow strengthening—told a different story. One plausible contributor to that gap is an AI-mediated consensus that overweighted certain narratives relative to the harder data.

There is an important difference in degree. Software monocultures create technical cascade risk you can patch. Interpretive monocultures create cognitive cascade risk you can’t—there’s no CVE number for a shared blind spot in judgment. The latter is softer and harder to measure. But the mechanism is familiar: reduced diversity of independent judgment.

I use these models constantly. They compress research, surface patterns I’d have missed, and force clearer thinking when I use them well—Claude caught an inconsistency in a cash flow assumption last month that I’d read past twice on my own. That’s real. The danger isn’t the tool. The danger is treating the tool as the authoritative voice rather than one input among many. The edge increasingly belongs to people who combine the model’s speed with proprietary data, primary research, domain experience, and a willingness to hold non-consensus views. Those who simply outsource the interpretation may find themselves more correlated than they realize, and won’t know it until the moment it matters.

Diversity of thought was never free. It was always work.

I noticed myself skipping the work, just for a second, on an ordinary Tuesday. That’s usually how it starts.

Categories
AI

The Quiet Trade-offs of Open Weights

An open letter is circulating this week — Open Weights and American AI Leadership — signed by a broad coalition of companies arguing that downloadable model weights are essential to U.S. competitiveness, diffusion of capability, and even safety. It makes a strong case on access, competition, and sovereignty. It also nods, briefly, to the fact that once weights are released they pass beyond the original developer’s control.

What it doesn’t fully reckon with are two structural realities that follow from that release. Neither is an argument against open weights. Both are simply facts about what openness costs, and what it buys.

Two core limitations

First, control.
Once the weights leave the developer’s servers, the developer can no longer dictate how the model is used. System prompts, refusal training, monitoring, rate limits, rapid safety updates — none of it reaches an independent deployment. Users can strip safeguards, fine-tune for purposes the original team would never sanction, or run the model somewhere it was never meant to go. The letter acknowledges the loss of control. It doesn’t linger on what that means for ongoing safety governance.

Second, learning.
Closed, hosted models draw on a continuous stream of real usage — the queries people actually ask, the reasoning traces that result, the places the model fails or succeeds in the wild. As appropriate that exhaust can be sampled, reviewed, and fed back into improvement. Open weights running independently offer no such path. The developer has no visibility into how the model is being used at scale once it’s out the door. Improvement then falls to slower, thinner channels: community datasets, published evals, distillation from any parallel closed models the lab still runs, internal preference data. The high-volume, real-distribution signal is gone.

These two limitations travel together. The same openness that strips the developer’s control also strips its ability to learn from the model’s actual use.

Sovereignty flips the perspective

A parallel argument has been building around “sovereignty” — an enterprise or government’s ability to own its data, its fine-tuned weights, its compute, its proprietary edge. In this framing, open weights are a path to control, but for the user, not the developer. The organization downloads the model, adapts it inside its own environment — often air-gapped — and keeps whatever capability results private. What the lab surrenders in ongoing control, the institution gains in independence.

But the same move that delivers sovereignty deepens the learning problem. An organization running the model under genuine sovereignty keeps its queries, reasoning traces, and institutional knowledge inside its own walls, by design. None of that returns to the developer. The more high-value users — governments, defense, critical infrastructure, large enterprises — choose sovereign deployments, the thinner the real-world signal available to the labs training the next generation of models. Local fine-tuning can still happen, but that learning stays private. It doesn’t flow back into the shared base model.

What the letter leaves out

The letter is right that closed models aren’t automatically safer, that concentration creates single points of failure, and that transparency invites broader scrutiny. It’s also right that open weights expand access and cut lock-in. Those points hold.

But it treats the developer’s loss of control mainly as a manageable risk that community examination can offset. It celebrates user control and sovereignty without mapping the full exchange: the developer loses both control and its richest usage signal, and that signal thins further as more institutions choose real sovereignty. The information environment models improve in is changed by these choices — not just the distribution of access.

Other distinctions worth naming

  • Update velocity. Closed models patch globally and immediately. Open-weight deployments lag; many users never leave an old version.
  • Customization power. The flip side of lost control is real specialization — downstream users can adapt a model far deeper into a narrow domain than its original developer ever will.
  • Transparency versus opacity. Open weights let outside researchers inspect and red-team a model in ways closed systems don’t allow.
  • Economic structure. Open weights commoditize the base model and push value toward data, fine-tuning, infrastructure, and applications.
  • Privacy at the edge. Running a model fully offline or on private infrastructure is a guarantee hosted services simply can’t match.

A clearer accounting

Open weights aren’t a free lunch. They’re a deliberate trade: the developer gives up ongoing control and the continuous signal of real usage, in exchange for diffusion, customization, outside scrutiny, and user independence. Institutional sovereignty amplifies one side of that trade — it solves the dependency problem for the user while further starving the developer of high-stakes, real-world feedback.

That trade may still be the right one for research progress, economic diffusion, spreading capability beyond a handful of labs, privacy-preserving deployment. But it’s a trade with real, compounding costs. Treating the loss of control as a footnote, and the loss of the learning signal as invisible, leaves an incomplete map.

The letter is right that American leadership will be judged by the strength of the whole ecosystem, not by any single frontier model. An accurate map of that ecosystem has to include what openness and sovereignty actually cost the original developers, in control and in learning both. Only then can we reason clearly about when those costs are worth paying — and what might offset them.

The conversation is better when we name the full set of trade-offs instead of talking around them.

Categories
AI China

Model Competition

The Chinese AI company Moonshot AI released its Kimi K3 model yesterday. I played with it a bit in the Kimi iOS app and, frankly, was underwhelmed. It was slow and disappointing. But this was hardly an extensive test. It’s strictly my first impression. I’m not doing coding. I’m what model makers would call a knowledge worker.

This morning Gavin Baker posted on X about Kimi K3 calling it potentially an “important inflection point for AI”:

Kimi K3 may be an important inflection point for AI. Potentially negative for Anthropic and OpenAI while being net positive for essentially every other company in the world. I mean that very literally. Although the real “Sputnik moment” would be an open-source frontier model that was also token efficient unlike Kimi K3 which is 50-70% more expensive to run than GPT 5.6 per Artificial Analysis.

Rationale:
 
A world where there are only 2-3 dominant frontier labs with 90% inference margins is net negative for every other layer while being awesome for those 2-3 labs. Those labs would become monopsonies for power, data centers, semiconductors and hyperscalers and would obviously vertically integrate over time into all those layers while also completely subsuming the application/software layers. 
 
Anything that lowers margins and increases competition at the model layer is good for every other AI layer: power, semiconductors, hyperscalers, neoclouds and yes even software.
 
This is why Jensen is so supportive of open-source. An open-source model requires the exact same amount of compute to run as a closed frontier model of similar size and architecture. Kimi K3 is roughly the same price as GPT 5.6 Terra on a per token basis, which actually suggests that it is less computationally efficient as I am sure that GPT 5.6 is priced to a higher margin than K3. And given that K3 is a token wastrel, i.e. token inefficient, it is significantly more expensive per task than GPT 5.6 and Grok 4.5, which are much more token efficient. Cost per token and token efficiency (i.e. intelligence density per token) are the drivers of intelligence per unit of cost. The winning AI companies will be those that offer the most intelligence per $ over time.
 
Lower margin % at the model layer = more margin $ at every part of the infrastructure layer and is a godsend for software. This can happen either through open-source models like K3 at the frontier or having a vertically integrated model company like Meta, SpaceX or Google at the frontier. Both outcomes result in a lower margin % at the model layer as vertically integrated model companies don’t really care where the margin $ come from. This is why it was so painful for OpenAI and Anthropic when Google was right there with them from a model competitiveness perspective and why Grok 4.5 and Muse 1.1 were just as important as Kimi K3.

The reason Kimi K3 is only potentially negative for Anthropic and OpenAI is 1) the @ericvishria point that the Claude and ChatGPT products and harnesses may be more important than their models today and 2) the hypothesis that they have much more advanced model checkpoints internally that are already being used for RSI. In the latter scenario, reaching RSI even a few months ahead of other labs might be enough to cement a permanent lead.

Time will tell on both points. And likely fairly quickly.

Caveat would be that since Kimi K3 is not token efficient and thereby actually more expensive than ChatGPT 5.6, we may need to see a more token efficient open-source model at the frontier or see Grok 5/Composer 4/Muse 2 at multiple points on the Pareto frontier for this potential risk to Anthropic and OpenAI to play out. And I am sure they will both vertically integrate as quickly as possible while continuing the product/harness strength they have shown over the last 8 months.

I wanted to explore this further so I asked Grok to explain his post:

This is a tweet (X post) from Gavin Baker (@GavinSBaker), Managing Partner and CIO of Atreides Management, posted on July 17, 2026.

It analyzes the recent launch of Kimi K3 (by Chinese AI company Moonshot AI) as a potentially significant shift in the AI industry. Baker argues it could hurt the high-margin dominance of closed U.S. frontier labs like OpenAI and Anthropic while benefiting nearly everyone else in the broader AI ecosystem (infrastructure, chips, data centers, power, hyperscalers, and software).

Quick Context on Kimi K3

Kimi K3 is Moonshot AI’s new flagship model (launched ~July 16, 2026). Key specs include:

  • ~2.8 trillion parameter Mixture-of-Experts (MoE) architecture.
  • Massive 1 million token context window.
  • Native vision capabilities.
  • Strong early performance in coding, agentic tasks, and long-context reasoning (competitive with or close to top closed models like GPT-5.6 variants or Claude models in certain benchmarks).
  • Pricing: $3 input / $15 output per million tokens (with caching discounts).
  • Important: Full weights are scheduled for open release (modified MIT license) around July 27, 2026 — making it one of the first truly open frontier-scale models.

It represents a step up from prior Kimi models and adds meaningful competition from a non-U.S. lab.1

The Two Images in the Tweet

These are screenshots from Artificial Analysis (an independent benchmarking site):

  1. Bar chart — “Cost per Intelligence Index Task” (lower is better). It breaks down costs by token type (input, output, reasoning, cache, etc.) and shows Kimi K3 as relatively expensive compared to some GPT-5.6 variants.
  2. Scatter plot — Intelligence Index (y-axis) vs. Cost per task (x-axis, log scale). It positions models in quadrants, with the “most attractive” area being high intelligence at low cost. Kimi K3 appears competitive on intelligence but not the cheapest.

These visuals support Baker’s point that while Kimi K3 reaches frontier-level intelligence, it is not the most token-efficient, making its effective cost per useful output higher than more efficient models.

Baker’s Core Argument (Broken Down)

Baker’s thesis is economic and structural, not just technical:

  • A world with only 2–3 dominant closed frontier labs (with ~90% inference margins) is bad for almost everyone else.
    Those labs would act as monopsonies (dominant buyers) for power, chips, data centers, and compute. They would vertically integrate upward (into infrastructure) and downward (into apps/software), capturing most of the value. Competition at the model layer prevents this concentration.
  • Anything that increases competition or lowers margins at the model layer is broadly positive.
    It redistributes economic value to the rest of the stack: more demand for GPUs/chips (inference especially), power, data centers, hyperscalers/neoclouds, and software (because cheaper/better AI drives wider adoption).
  • Token efficiency matters more than raw price per token.
    Baker notes Kimi K3 is roughly the same per-token price as GPT-5.6 but 50-70% more expensive overall according to Artificial Analysis. It is a “token wastrel” — it uses more tokens to achieve similar results (lower “intelligence density per token”).
    → True winners deliver the most intelligence per dollar. Cost-per-token + efficiency together determine real value.
  • Why NVIDIA’s Jensen Huang loves open-source models.
    Open models still require massive compute to run at scale (same as closed ones). More models + wider adoption = more inference demand = more GPUs sold. A few ultra-high-margin closed labs would concentrate spend and reduce overall volume.
  • Vertical integration changes the game.
    Companies like Google, Meta, or xAI/SpaceX that have both strong models and their own infrastructure don’t obsess over model-layer margins — they capture value elsewhere. This is why more competition (open or vertically integrated) hurts pure-play high-margin model companies like OpenAI/Anthropic.
Why It’s Only “Potentially” Negative for OpenAI and Anthropic

Baker is careful:

  • Their products, user interfaces, and “harnesses” (the surrounding tools/ecosystems) may matter more than raw model performance right now.
  • They likely have more advanced internal models/checkpoints (possibly already used for recursive self-improvement/RSI).
  • Kimi K3’s token inefficiency softens the blow for now. A more efficient open-source frontier model (or stronger showings from Grok, etc.) would be more threatening.
  • OpenAI and Anthropic are already moving to vertically integrate.
Overall Takeaway from the Post

Baker sees Kimi K3 (and the wave of strong Chinese models like Qwen, DeepSeek, GLM) as healthy competition that prevents unhealthy concentration. It’s not yet a full “Sputnik moment” (a truly dominant, efficient open-source frontier model), but it points in that direction.

Net effect he describes: Lower % margins at the model layer → more total margin dollars flowing to infrastructure, chips, power, and software layers. This is bullish for the broader AI ecosystem even if it pressures the pure closed frontier labs.

The tone is analytical and balanced — he acknowledges uncertainties and that time will tell (likely quickly). Replies to the post generally agree this kind of competition is positive for the industry overall and could shift more spend toward efficient inference hardware.

This fits Baker’s broader investing views on AI: he often emphasizes infrastructure winners, token economics, and how value accrues across the full stack rather than just at the smartest model.

Interesting.

Categories
AI Podcasts

A Remarkable Conversation…

Highly recommend this conversation between Harry Stebbings and Clay Bavor. Among many topics, I especially enjoyed the discussion about not investing in frontier models, the important values, the particular importance of craftsmanship, intensity, and family. And the special conversation about parenting and kids near the end. Just a delightful conversation to be able to enjoy!

Key Highlights:

• Founding Sierra: Bavor explains why he and Taylor chose to start Sierra, focusing on the transformative potential of language model-based agents (1:37 – 5:53).
• The AI Tech Stack: Sierra focuses on building enterprise-grade agent architectures and fine-tuning models on top of open-weights models rather than pre-training foundation models from scratch, prioritizing capital efficiency (5:53 – 7:15).
• Unbounded Demand for Intelligence: Bavor argues that there is massive, unmet demand for “frontier-level” intelligence in fields like coding, science, and legal work (7:15 – 11:41).
• Internal AI Operations: He details the use of Pinecone, an internal AI agent Sierra developed to navigate company data, streamline engineering, and assist in recruitment (18:36 – 22:00).
• Enterprise Strategy: Sierra employs a “forward-deployed” engineering model, embedding staff within client companies to ensure rapid, effective integration of AI, leading to quick deployment timelines (30:12 – 33:22).
• Board Governance: To keep pace with the speed of AI development, Sierra operates on a six-week board meeting cadence, utilizing comprehensive memos instead of traditional slide decks (39:07 – 41:13).
• Corporate Culture: Bavor emphasizes values like craftsmanship, intensity, and family. He also highlights the importance of working in-person to foster apprenticeship, mentorship, and a cohesive team culture (43:02 – 55:41).

Categories
AI AI: Large Language Models Apple

The Slipstream Strategy

Apple had a problem no amount of money could solve. An iPhone can’t draw the power or shed the heat of a data center, so ten different tasks can’t mean ten different models fighting for the same sliver of RAM. Apple’s answer was to freeze one small, efficient base model into the device and then swap tiny adapters in and out of it in milliseconds — a summarization adapter for your texts, a Siri adapter for on-screen actions, and a handoff to Private Cloud Compute for anything heavier. The phone behaves like it’s running many models. It’s running one model wearing many hats.

That architecture — a frozen base plus swappable adapters — is quietly becoming the default way serious AI companies build, and it’s worth understanding why, because it inverts the assumption most people still carry into this industry.

The assumption is that winning means owning a frontier model. Sierra co-founder Clay Bavor pushed back on that on a recent 20VC episode: pouring capital into your own pre-training, he argued, tends to leave you holding a highly perishable bag of floating-point numbers. Open-weight models improve fast enough that yesterday’s frontier is next quarter’s commodity. The companies playing this well aren’t racing to out-spend the labs. They’re slipstreaming behind them — taking the free, state-of-the-art engine and putting all their effort into what sits on top of it.

What sits on top is LoRA — low-rank adaptation. The old failure mode was catastrophic forgetting: fine-tune a model hard enough on your own data and it forgets how to reason generally. LoRA sidesteps this by leaving the base model untouched and training a small set of additional parameters alongside it — a thin layer of expertise bolted onto a frozen foundation. You get real domain depth without touching the thing that makes the model work at all.

The business logic that follows from this is the actual point, and it’s simpler than it looks:

You stop being hostage to any one model provider — if a better open-weight model ships next month, you port your adapter, not your whole product. You can serve hundreds of differently-customized clients off one base model on one piece of hardware, instead of running a separate giant model per customer. You can ship a fix in an afternoon, because an adapter is a few hundred megabytes, not a training run. And in regulated industries, your proprietary data can train an adapter that never leaves your own infrastructure.

None of this is really a story about model architecture. It’s a story about where the moat moved. For a while the moat was raw capability — whoever had the best model won. Apple and Sierra are betting the moat is now somewhere else entirely: in how tightly you can weave a commodity intelligence into a specific workflow, a specific dataset, a specific customer relationship. The engine is free. The adapter is the business.

Categories
AI Apple Google

The Floor

I compared the frontier to a three-star chef making grilled cheese in “Context Rot” — the smartest models on earth spending most of their time on work beneath them, the way a chef trained at Le Bernardin might still melt cheese between two slices of bread on a Tuesday night and call it dinner. The comfort was the point: if the sharpest tool is saved for hard problems and something merely-very-good handles the rest, nobody’s losing anything. The floor was never the interesting part.

I’ve kept turning the joke over, and I think I had the wrong worry.

Watch what companies do with their AI spend, not what they say. Coinbase moved engineers off frontier models onto open weights and cut its AI spend nearly in half while usage kept climbing. Nvidia runs a closed model as orchestrator and routes the actual volume — the daily uncelebrated bulk of it — to open weights it controls. The frontier is becoming a dispatcher, deciding where the request goes and rarely doing the work itself. The instinct is to worry about whose open weights end up running that volume, and right now the most capable ones at scale are Chinese — GLM, Kimi — which makes it tempting to read this as a contest America is quietly losing: the floor of the AI economy built somewhere else, at a price export controls can’t touch. You cannot embargo a file already downloaded. You cannot price-match free.

But that framing has a hole. Google’s own Gemma family is open-weight and good enough to handle that daily volume without anyone reaching for GLM or Kimi. “Open weights are a Chinese story” only holds if you don’t count the open models the company running Android and half the internet’s search traffic has already shipped.

And once I saw that hole, a bigger one opened behind it. I’ve been trying Apple’s new Siri — arriving with iOS 27 this fall, genuinely surprisingly good in beta — and it made me realize open weights, of any nationality, were never going to cook most of the world’s dinners. Apple and Google are.

Consider what actually determines where the world’s routine inference runs. Not which model benchmarks best, not which weights are downloadable — what’s already installed. Apple ships to well over a billion active devices before routing a single query through Siri’s new architecture. Nobody has to be persuaded to try it, or hear about it on a podcast; it’s the thing that answers when you press the button you’ve pressed for a decade. Google owns the search bar and the Android default the same way. Between them, that’s most of the world’s phones — and phones are where most of the world’s questions get asked.

The open-weight framing assumes the floor is up for grabs, that whoever ships the best free model wins the daily grind by merit. But the floor was never a bazaar. It’s a set of defaults, owned by whoever already has the device in your hand, not whoever holds the most generous license. Apple didn’t need to win the model war to win this. Its heaviest reasoning tier is built with Google, running on Nvidia chips in Google’s cloud, under a deal reported at roughly a billion dollars a year — Apple doesn’t fully own the engine doing the thinking. It doesn’t need to. It owns the button.

That’s a quieter concentration than an export-controls fight, and a harder one to dislodge. An open model can be forked, distilled, undercut, or out-competed by the next release. A billion phones with an assistant built into the lock screen cannot be routed around. Whoever’s weights hum underneath barely matters, the way it barely matters to a diner which supplier delivered the flour. What matters is whose kitchen the meal came from, and whose name is on the door.

The grilled-cheese chef was never the risk. Two chefs are about to own nearly every kitchen on earth, and most of us will never notice — because a kitchen you’ve been eating out of for a decade doesn’t feel like something that was won. It just feels like home.

Owning the kitchen and getting paid for what’s cooked in it, though, turn out to be two different questions. That one’s for another post.

Categories
AI Consulting

The Judgment Layer

An analyst’s note about the CEO of one of the largest consulting companies making comments at an investor conference includes a line that deserves more attention than it got: “token volume used on a project isn’t a proxy for AI maturity.”

Translation — clients are burning money on frontier models for problems that don’t need frontier models, and they’re not getting the outcomes they expected.

This firm’s CEO offered this as a business opportunity. I read it as a confession.

The old consulting model was simple: client has a technology problem, firm deploys humans to solve it. Billing followed effort. The new problem is different in kind — clients have an AI strategy problem. They know they’re supposed to be using AI. They’ve heard the word “frontier.” They’re spending accordingly. They just don’t know why, and the outcomes are showing it.

So the CEO is right that there’s an opportunity here. The value proposition shifts from implementation to judgment — not deploying AI, but knowing when not to deploy the expensive one. Matching capability to problem. Being trusted enough to tell a client that their $50M frontier model contract is solving a $500K problem.

Here’s the irony that the comment skates past: that advice is structurally difficult for a large consultancy to give.

The business model that built consulting firms was billing for doing. The more you deploy, the more you bill. Helping a client spend less, or choose the cheaper model, or run a narrower project, is genuinely good advice that the incentive structure actively works against. You don’t grow a $70 billion professional services firm by talking clients out of scope.

The judgment layer, if it becomes the real value, requires something closer to a doctor’s relationship with a patient than a contractor’s relationship with a client. Doctors get paid whether they prescribe or not. The value of the visit is the diagnosis — including the diagnosis that says you don’t need the expensive intervention. Consultants, historically, get paid to prescribe, and paid more when the prescription is larger.

There’s a reason we trust doctors with that asymmetry and not contractors. Licensing, malpractice, professional norms built over centuries — all of it exists to align the incentive. Consulting has none of that infrastructure. What it has instead is reputation, which is slower-acting and easier to game.

Whether the large firms can actually make the shift — rather than just reframe the same billable-hours model in the language of AI optimization — is the real question the market is wrestling with. The CEO’s comment is genuinely perceptive about where client value lies. It’s less clear that consulting firms are currently built to capture it honestly.