Categories
AI Learning Meta

Trajectories: how Meta plans to make Muse smarter by watching it work

Buried in the data policy section of Meta’s long post on how it built safety into Muse is a sentence that isn’t about safety at all:

“Inference data, the back and forth conversations between you and your Muse and the tool calls and subagent handoffs that result (‘trajectories’) are useful data for training new checkpoints of the LLM model at the core.”

Two things are worth noticing. First, the technique described isn’t new — training on agent rollouts is standard practice across the field. Second, the company describing it is Meta, in plain language, in a public post, about a product aimed at billions of consumers. The labs usually discuss this stuff in papers about coding agents. Meta just told its future user base: your agent’s work product is our training data. The candor is the story, not the technique.

A trajectory isn’t a chat log. It’s the complete record of an agent doing a job: what you asked, what it tried, which tools it called, where it went wrong, how it recovered, which subagents it spawned, and whether the thing actually got done. Every time you let Muse book the flight, triage the inbox, or research the supplier, you’re generating one.

From text to behavior

The technique matters anyway, because the diet that AI trains on is changing. Pretraining was about text — the whole internet, more or less. Post-training was about preferences — which answer humans liked better. Trajectories are the third course: demonstrations of competent behavior, in full, mistakes included.

There’s a reason for the shift. Text teaches a model what the world looks like. Preferences teach it what people want. But neither teaches it how to do a 40-step task without wandering off, recovering from a dead end, or knowing when to ask for help. That only exists in records of agents actually doing things. And until recently, almost nobody had those records at scale — because almost nobody had agents doing real work at scale.

The demonstrated instance — and what it doesn’t prove

Meta’s concrete example is Muse Spark 1.2, co-trained with Muse Code: model and harness trained together on rejection-sampled harness trajectories — run the agent many times, keep the runs that succeeded, train on those — with recipe-level tuning for goals, context compaction, and subagents. The model isn’t learning to predict text; it’s learning to behave inside a specific set of tools.

This is the end of the “base model plus clever prompting” era. The artifact is the bundle — model and harness, co-designed. A model trained on trajectories from one harness will be genuinely better inside that harness than a smarter general model dropped into it cold. Meta is saying this out loud; OpenAI and Anthropic are doing the same thing more quietly.

But notice the domain: coding. And coding is exactly where trajectories are cheapest to manufacture — verifiable unit tests, sandboxed repos, SWE-bench-style tasks. Nothing about the Muse Code result requires a single consumer or a single inbox. So the one demonstrated instance of Meta’s trajectory training sits squarely in the category where Meta’s distribution advantage matters least. Meta hasn’t shown its hand on the category that actually matters.

The other category is the personal one, and there the evidence is thinner. What exists is a stated intent, not a published result. The data policy says personal Muse trajectories “are useful data for training new checkpoints.” The product is designed to generate them: Meta’s own design example has Muse monitoring school emails, adding dates to a family calendar, filling a supply cart, finding a sale sweatshirt, booking dinner, and catching a sports tryout deadline hours before it closed. That is what an unverifiable-domain trajectory looks like — a morning of small judgments no unit test could grade.

No training run on that data has been published. No benchmark, no “Muse got X% better at inbox triage after training on Y million user trajectories.” So the sharpest version of the argument — that the real moat is the data nobody else can fake — should be labeled for what it is: a prediction, not an observed fact. It’s a prediction with a mechanism, though: these are judgments that can’t be synthesized, in the one distribution channel that reaches the people making them.

The flywheel — and its limits

With that caveat on the table: trajectories get better with scale, and Meta has scale like nobody else: billions of users across its apps, and now an agent — Muse — sitting inside them. Every user interaction is a potential training trajectory. Better trajectories train a better model; a better model makes a better agent; a better agent attracts more users. Meta states the bargain plainly: “every Muse user gets a better personal agent as we all collectively use the product and help the model understand the intricacies of human life.”

But “most users = most trajectories = structural advantage” needs its counter-case, because a lot of the highest-value trajectory data right now doesn’t come from consumers at all. It comes from sandboxes, the same kind that produced Muse Code. Synthetic and simulated trajectories sidestep the need for billions of users entirely — Anthropic and OpenAI are getting rich trajectory data from developers running Claude Code and Codex against real repos, no social-app distribution required.

The honest version of the moat argument is narrower, and more interesting. Synthetic trajectories work brilliantly where success is verifiable — code either passes the tests or it doesn’t. They work poorly where success is a matter of judgment: triaging an inbox, planning a trip around someone’s actual preferences, knowing which email deserves a reply. There is no unit test for a life well managed. And those unverifiable, deeply personal tasks are exactly what Meta means by “personal superintelligence” — and exactly where its distribution gives it trajectories nobody else can synthesize. The moat isn’t “most data.” It’s “the data nobody else can fake.”

The price of the flywheel

There’s a wrinkle, and Meta knows it. The flywheel runs on your data — your emails, your calendar, the messy reality of your life, which is exactly what makes the trajectories valuable. Meta’s answer is sanitization (“trajectories are sanitized to remove key personally identifiable information”), an opt-out switch, no sharing with ad systems, and a forthcoming “Confidential VM” that would cryptographically prevent even Meta from seeing your data.

The tension is fundamental, and it’s the sharpest part of the whole picture: the product gets smarter by watching you, and it earns the right to watch you by being trustworthy. Those two imperatives pull in opposite directions, and no amount of engineering fully resolves it — the Confidential VM, if it ever ships as described, would resolve it by breaking the flywheel, since trajectories Meta can’t see are trajectories Meta can’t train on. The opt-out rate will be the market’s verdict on the deal Meta is offering.

Experience is the missing piece

But the deepest reason trajectories matter has nothing to do with Meta’s strategy. It’s about what intelligence actually is.

A model trained only on text knows the world the way a brilliant student knows it from books. A model trained on trajectories knows it the way a practitioner does — from doing the thing, failing at it, and adjusting. The trajectory is the closest thing AI has to experience. And an agent that records its experience, keeps what worked, and folds it back into itself is doing something that rhymes with learning.

This is why I keep coming back to continual learning as the critical missing piece in AI. The models are frozen at training time; everything they “learn” afterward lives in context windows and memory files, fragile and local. Trajectories are the bridge: today’s version of the loop is slow and centralized (collect trajectories, train a new checkpoint, ship it), but the direction is obvious. The end state is an agent that learns continuously from its own experience — from your experience with it — the way people do.

Meta’s bet is that the path to personal superintelligence runs through watching agents work, at planetary scale, and distilling what works back into the model. No result yet proves the bet pays off — the personal trajectories are still a hypothesis, not a track record. But it’s an unglamorous hypothesis, no new scaling law, just better data about doing things, and unglamorous bets about data have a good track record in this field. The internet made the last generation of models. Trajectories might make the next one — if Meta can show, and not just say, that the data nobody else can fake is data that actually teaches.

Categories
AI Anthropic Apple Google OpenAI

It’s the Harness, Stupid!

I’ve been wondering whether we’ve been looking at the AI stack from the wrong end.

Recently Kris Patel on X laid out a set of excellent questions that he’s looking to have answered as Anthropic and OpenAI move toward going public:

  1. Do you really need frontier-scale intelligence for every task?
  2. Can open-weight models provide an effective alternative to frontier models at a significant discount?
  3. Is the ultimate moat the intelligence or the harness?
  4. What other business models will the frontier labs have to adopt to make the unit economics work long term?
  5. How are you going to prevent distillation from capturing your IP and releasing it?

I’m going to explore only the third question here — moat versus harness. The other four deserve their own consideration, particularly once we have the Anthropic and OpenAI S-1s in hand, revealing for the first time the unit economics of the two largest frontier labs and how much runway they have to support the capacity they’ve contracted.

By “harness” we mean everything surrounding the model: the interface, context, memory, tools, orchestration, evaluation, permissions, and increasingly the user’s accumulated habits and data. The model supplies intelligence. The harness turns intelligence into a product.

Listening to Gavin Baker on the recent All-In episode sharpened this line of thought into something more concrete. He referenced a thought experiment from Eric Vishria: even if OpenAI or Anthropic lost their edge at the pure model layer, they would still retain significant value because of the product harness—the interface, the surrounding tooling, the orchestration—and the user familiarity and habits that have already formed around those platforms. Baker said there is a strong element of truth to it. I think he’s right, and the reasoning behind it is worth spelling out. A model can be replicated, distilled, open-weighted, or commoditized. A mature harness has network effects, switching costs, proprietary context, distribution, workflow integration, and accumulated user behavior. That’s a much harder thing to dislodge.

We are watching intelligence become more abundant and more interchangeable at the same time that the systems built around that intelligence are becoming stickier. On the developer side, the strongest examples are already clear. Cursor turns the IDE into a multi-model agentic environment with deep codebase awareness. Claude Code runs long-horizon coding agents from the terminal, planning, editing, testing, and iterating. Grok Build, Claude Cowork and similar tools emphasize parallel agents and tighter control over local context. In each case the model is a component; the surrounding system does the real work of routing, memory, tool use, and evaluation.

The consumer version of the same idea is now taking clearer shape at Apple. The rebuilt Siri AI shown at WWDC 2026 is not trying to win the pure model race. It is built as a personal harness. A system orchestrator decides what stays on-device with Apple’s Foundation Models, what moves to Private Cloud Compute, and when heavier reasoning is required. Personal context—messages, email, photos, calendar, notes, on-screen awareness—is handled largely on-device through the Spotlight semantic index and App Toolbox. Apple is designing the system so that personal context can be used without giving Apple itself access to it. Conversation history lives in a dedicated Siri app for the user to revisit. And it all syncs across all your Apple devices.

Here is the part I think matters most, and it’s easy to miss if you only read the privacy story. Apple’s Foundation Models framework doesn’t just call Apple’s own models—it’s built to support cloud models from other providers, including Claude and Gemini, conforming to a common protocol. Which means the system orchestrator, not the user, decides which model handles which task. This request goes to the on-device model. That one goes to Private Cloud Compute. A harder one might go to Claude or Gemini. The user doesn’t need to choose, and increasingly doesn’t need to know.

That’s the inversion worth exploring further. The frontier model stops being the interface and becomes a component underneath someone else’s interface. The harness chooses the intelligence. And the company that owns the harness—the OS, the identity layer, the permissions, the apps, the sensors, the notifications, the semantic index tying all of it together—has a form of leverage that has very little to do with whose model is smartest this quarter.

That reframes the subscription question too. I don’t think the right question is whether Siri gets good enough to beat ChatGPT or Claude at reasoning. I think Siri doesn’t need to win that fight at all. It needs to win a different layer entirely—the ambient assistant layer, not the reasoning layer. They’re doing different tasks. ChatGPT or Claude might remain where you go when you think, when I need to reason about something. Siri becomes where I go when I need something done: find (or make) my reservation, text my friend, find that old photograph, update my shopping list, schedule that meeting, add this thought to my notes, figure out when we’re free next week, remind me about that thing we discussed three months ago. Apple’s advantage as an ambient assistant isn’t primarily that it has your personal data. It’s that it has OS-level authority over the world that my personal data lives in.

Of course this is still early. Execution will determine how much of the architectural promise becomes daily reality. Reliability, agentic follow-through, and the quality of the on-device models will matter as much as the privacy story or the multi-model routing. But the strategic bet itself is clear, and it aligns with the broader shift: durable value is migrating toward the systems built around the models, especially systems that sit atop private, permissioned, personal context that competitors cannot easily reach. My early personal experience with the new Siri in iOS 27 betas has impressed me so far. All of this also seems to apply to Google in the context of their Pixel family of devices.

This doesn’t mean frontier labs lose. Pricing power still exists at the high end for the hardest agentic and long-horizon work. Open-weight models will continue to pressure costs and expand access. Distillation remains a real risk. But the more the capability gap narrows, and the more a harness like Apple’s can route among interchangeable frontier models rather than depend on any single one, the stronger the case that value settles into whoever controls the context—not whoever trained the model.

The coming Anthropic and OpenAI S-1s will tell us whether the frontier labs can make their economics of intelligence work. The next generation of Siri, Gemini, ChatGPT, Claude, and whatever comes after them may tell us something even more important: who gets to own the primary relationship with the user.

The model may be the engine. But the harness is where the driver sits.

What a time to be alive!