Categories
AI Learning Meta

Trajectories: how Meta plans to make Muse smarter by watching it work

Buried in the data policy section of Meta’s long post on how it built safety into Muse is a sentence that isn’t about safety at all:

“Inference data, the back and forth conversations between you and your Muse and the tool calls and subagent handoffs that result (‘trajectories’) are useful data for training new checkpoints of the LLM model at the core.”

Two things are worth noticing. First, the technique described isn’t new — training on agent rollouts is standard practice across the field. Second, the company describing it is Meta, in plain language, in a public post, about a product aimed at billions of consumers. The labs usually discuss this stuff in papers about coding agents. Meta just told its future user base: your agent’s work product is our training data. The candor is the story, not the technique.

A trajectory isn’t a chat log. It’s the complete record of an agent doing a job: what you asked, what it tried, which tools it called, where it went wrong, how it recovered, which subagents it spawned, and whether the thing actually got done. Every time you let Muse book the flight, triage the inbox, or research the supplier, you’re generating one.

From text to behavior

The technique matters anyway, because the diet that AI trains on is changing. Pretraining was about text — the whole internet, more or less. Post-training was about preferences — which answer humans liked better. Trajectories are the third course: demonstrations of competent behavior, in full, mistakes included.

There’s a reason for the shift. Text teaches a model what the world looks like. Preferences teach it what people want. But neither teaches it how to do a 40-step task without wandering off, recovering from a dead end, or knowing when to ask for help. That only exists in records of agents actually doing things. And until recently, almost nobody had those records at scale — because almost nobody had agents doing real work at scale.

The demonstrated instance — and what it doesn’t prove

Meta’s concrete example is Muse Spark 1.2, co-trained with Muse Code: model and harness trained together on rejection-sampled harness trajectories — run the agent many times, keep the runs that succeeded, train on those — with recipe-level tuning for goals, context compaction, and subagents. The model isn’t learning to predict text; it’s learning to behave inside a specific set of tools.

This is the end of the “base model plus clever prompting” era. The artifact is the bundle — model and harness, co-designed. A model trained on trajectories from one harness will be genuinely better inside that harness than a smarter general model dropped into it cold. Meta is saying this out loud; OpenAI and Anthropic are doing the same thing more quietly.

But notice the domain: coding. And coding is exactly where trajectories are cheapest to manufacture — verifiable unit tests, sandboxed repos, SWE-bench-style tasks. Nothing about the Muse Code result requires a single consumer or a single inbox. So the one demonstrated instance of Meta’s trajectory training sits squarely in the category where Meta’s distribution advantage matters least. Meta hasn’t shown its hand on the category that actually matters.

The other category is the personal one, and there the evidence is thinner. What exists is a stated intent, not a published result. The data policy says personal Muse trajectories “are useful data for training new checkpoints.” The product is designed to generate them: Meta’s own design example has Muse monitoring school emails, adding dates to a family calendar, filling a supply cart, finding a sale sweatshirt, booking dinner, and catching a sports tryout deadline hours before it closed. That is what an unverifiable-domain trajectory looks like — a morning of small judgments no unit test could grade.

No training run on that data has been published. No benchmark, no “Muse got X% better at inbox triage after training on Y million user trajectories.” So the sharpest version of the argument — that the real moat is the data nobody else can fake — should be labeled for what it is: a prediction, not an observed fact. It’s a prediction with a mechanism, though: these are judgments that can’t be synthesized, in the one distribution channel that reaches the people making them.

The flywheel — and its limits

With that caveat on the table: trajectories get better with scale, and Meta has scale like nobody else: billions of users across its apps, and now an agent — Muse — sitting inside them. Every user interaction is a potential training trajectory. Better trajectories train a better model; a better model makes a better agent; a better agent attracts more users. Meta states the bargain plainly: “every Muse user gets a better personal agent as we all collectively use the product and help the model understand the intricacies of human life.”

But “most users = most trajectories = structural advantage” needs its counter-case, because a lot of the highest-value trajectory data right now doesn’t come from consumers at all. It comes from sandboxes, the same kind that produced Muse Code. Synthetic and simulated trajectories sidestep the need for billions of users entirely — Anthropic and OpenAI are getting rich trajectory data from developers running Claude Code and Codex against real repos, no social-app distribution required.

The honest version of the moat argument is narrower, and more interesting. Synthetic trajectories work brilliantly where success is verifiable — code either passes the tests or it doesn’t. They work poorly where success is a matter of judgment: triaging an inbox, planning a trip around someone’s actual preferences, knowing which email deserves a reply. There is no unit test for a life well managed. And those unverifiable, deeply personal tasks are exactly what Meta means by “personal superintelligence” — and exactly where its distribution gives it trajectories nobody else can synthesize. The moat isn’t “most data.” It’s “the data nobody else can fake.”

The price of the flywheel

There’s a wrinkle, and Meta knows it. The flywheel runs on your data — your emails, your calendar, the messy reality of your life, which is exactly what makes the trajectories valuable. Meta’s answer is sanitization (“trajectories are sanitized to remove key personally identifiable information”), an opt-out switch, no sharing with ad systems, and a forthcoming “Confidential VM” that would cryptographically prevent even Meta from seeing your data.

The tension is fundamental, and it’s the sharpest part of the whole picture: the product gets smarter by watching you, and it earns the right to watch you by being trustworthy. Those two imperatives pull in opposite directions, and no amount of engineering fully resolves it — the Confidential VM, if it ever ships as described, would resolve it by breaking the flywheel, since trajectories Meta can’t see are trajectories Meta can’t train on. The opt-out rate will be the market’s verdict on the deal Meta is offering.

Experience is the missing piece

But the deepest reason trajectories matter has nothing to do with Meta’s strategy. It’s about what intelligence actually is.

A model trained only on text knows the world the way a brilliant student knows it from books. A model trained on trajectories knows it the way a practitioner does — from doing the thing, failing at it, and adjusting. The trajectory is the closest thing AI has to experience. And an agent that records its experience, keeps what worked, and folds it back into itself is doing something that rhymes with learning.

This is why I keep coming back to continual learning as the critical missing piece in AI. The models are frozen at training time; everything they “learn” afterward lives in context windows and memory files, fragile and local. Trajectories are the bridge: today’s version of the loop is slow and centralized (collect trajectories, train a new checkpoint, ship it), but the direction is obvious. The end state is an agent that learns continuously from its own experience — from your experience with it — the way people do.

Meta’s bet is that the path to personal superintelligence runs through watching agents work, at planetary scale, and distilling what works back into the model. No result yet proves the bet pays off — the personal trajectories are still a hypothesis, not a track record. But it’s an unglamorous hypothesis, no new scaling law, just better data about doing things, and unglamorous bets about data have a good track record in this field. The internet made the last generation of models. Trajectories might make the next one — if Meta can show, and not just say, that the data nobody else can fake is data that actually teaches.

Categories
AI Business

The Wage of Knowing

In 1973 the Los Angeles Public Library installed a telephone line that worked while the building was dark. Dial H-O-O-T-O-W-L on a rotary phone, nine at night until one in the morning, and a librarian would answer. Somebody wanted to know the boiling point of mercury, or who wrote a poem they half remembered, or how many wives Henry VIII actually had, and a person on the other end of a cord found out. This went on for years. Nobody thought of it as data collection. It was just a service, a courtesy, a woman at a desk with a card catalog in her head.

I worked, in another life, in the payments industry, back when a merchant who wanted to charge your card had to call in and ask permission. There were rooms for this. Banks of phones, a bulletin of stolen numbers updated by hand, a floor limit past which a supervisor had to be found. The people answering the phones were, more often than you would guess, college students. Twenty years old, minimum wage, deciding in real time whether a stranger’s card was good. Nobody trained them for six months first. They learned the bulletin, they learned to listen for something wrong in a voice, and they said yes or no.

I have been driven, recently, by a car with nobody driving it. I noticed the wheel turning on its own and I braced for the wrongness of it. Thirty seconds later I was not bracing. I was looking out the window. The data says I was right to relax: across two hundred and twenty million miles, the cars involved in this experiment cause a small fraction of the serious crashes a human would have caused over the same roads. I did not need the data. I needed thirty seconds.

None of these people knew what they were doing. That is the thing about the librarian and the college student and, for that matter, about me learning to trust a wheel that moves by itself. The librarian was not building a search engine. The clerk was not training a fraud model. He was making rent. Their competence was not evidence, to them. It was just Tuesday. It became evidence later, to someone else, in a room they never saw — the accident logs, the chargeback data, the accumulated record of a million correct guesses that turned out to be exactly the material a system needed to learn the job and take it.

This is the part that is easy to get wrong. It is not that the human failed and the machine succeeded. It is that the human succeeding, over and over, in full view, was the demonstration that the job could be learned. You do not automate a task nobody can do. You automate the one being done well enough, often enough, for long enough that the pattern becomes visible. Doing the job right was never neutral. It was the case being built.

Which brings me to a woman I will call the lawyer, because there are thousands of her and none of them are exactly her. She has a laptop open at her kitchen table. She logs into a dashboard belonging to a company that pairs credentialed people with the AI labs that need them — a doctor here, a banker there, a corporate attorney with fifteen years of contract law behind her. She reads a model’s draft of a merger agreement and marks where it reasons like a first-year associate instead of a partner. She rewrites a clause. She explains, in the margin, why the model’s version would get laughed out of a negotiation. She is paid well for this. More, some weeks, than she billed certain clients.

She knows exactly what she is doing. That is the difference between her and the other three. The librarian did not know she was leaving a trail. The clerk did not know his good judgment would become someone else’s weights. I did not know, thirty seconds into that ride, that I was participating in anything at all. The lawyer knows. She is being paid, by the hour, at a rate that respects her expertise, to make her expertise legible enough that it no longer requires her. The company she works for has a name for this. They call it the reinforcement learning economy, which is a tidy way of saying: teach it everything, and then it will not need to call you back.

She does the work anyway. The rate is good. The work is interesting, in the way that teaching is interesting — you learn what you know by trying to say it clearly enough for someone else to use. Nobody is lying to her. The dashboard does not pretend to be anything other than what it is. She logs off at the end of the session the way anyone logs off after a long day of being excellent at something, tired in the specific way that comes from careful work, and she does not, from what I understand, spend the evening thinking about what she has just fed into the machine.

I keep coming back to the rotary dial. Somebody dialing H-O-O-T-O-W-L at midnight in 1973 could not have imagined the lawyer at her kitchen table. But the shape is the same, if you look at it long enough. A person answers a question well. The answering becomes a record. The record becomes a system. The system answers next time. Nobody in the room ever decided this was the plan. It just turned out, every time, to be the plan.

Categories
AI Creativity Living

Our Human Operating System: Is Our Upbringing Our Personal “System Card”?

Note: the following post was largely generated by Google Gemini 2.5 Flash. I prompted Gemini to draft it after reading Simon Willison’s post about the Claude 4 Opus system prompt and being struck by the notion of us humans also having our versions of system cards. I asked Gemini to probe and explore that notion along with the related notion of how our life experiences constitute the human version of reinforcement learning. Rather than avoid the use of and being critical of using AI to write for me, I’m enjoying exploring and learning more about its capabilities! One thing is clear: Gemini 2.5 Flash seems to be an impressive new model!

Simon Willison’s recent dive into the Claude 4 Opus system prompt got me thinking. He dissects the meticulously crafted instructions that define Claude’s core behavior, its ethical guardrails, and its fundamental operational parameters. It’s a fascinating glimpse into how a complex AI is given its foundational “personality” and purpose. But as I read, a parallel began to emerge in my mind, one that brought me back to something far more organic and familiar: ourselves.

Could it be that what we, as humans, are taught and learn from our parents and primary caregivers is, in essence, our own unique, individual “system card”?

Think about it. From the moment we are born, we are immersed in a world of instruction, observation, and subtle conditioning. Our parents, whether consciously or unconsciously, are constantly programming us. They instill values: “Always be kind,” “Honesty is the best policy.” They teach us social norms: “Say please and thank you,” “Don’t interrupt.” They guide our understanding of the world: “Look both ways before crossing,” “Stranger danger.” They impart their wisdom, their fears, their hopes, and their biases, all of which become foundational layers in our burgeoning minds.

This isn’t merely about rote memorization or factual knowledge. It’s about the deep-seated principles that govern our reactions, our decision-making, and our very perception of reality. Just as Claude’s system prompt dictates its default tone and its approach to difficult queries, our upbringing shapes our inherent optimism or pessimism, our tendency towards introversion or extroversion, our inclination to trust or to be cautious.

Consider the parallels more closely. A system prompt aims for consistency and predictability in an AI’s behavior. Similarly, parents strive to create a stable and predictable environment for their children, instilling routines and expectations that foster a sense of security and belonging. This consistency helps to solidify the early “programming.”

The “ethical guardrails” in an AI system prompt are designed to prevent harmful or undesirable outputs. Our parents, too, establish ethical guardrails. They teach us right from wrong, the consequences of our actions, and the importance of empathy. These lessons, often reinforced through discipline and encouragement, become our internal compass, guiding us away from behaviors that could harm ourselves or others.

Furthermore, a system prompt often defines an AI’s learning parameters and its ability to adapt. Our upbringing isn’t a static, one-time download. It’s an ongoing process. As we grow, we continue to learn from our parents through their reactions to new situations, their advice on navigating challenges, and their own evolving perspectives. This continuous input refines and expands our internal “system card,” allowing us to adapt to new information and experiences.

Of course, the analogy isn’t perfect. We are not machines, and our development is infinitely more complex and nuanced than any AI’s. We possess free will, consciousness, and the capacity for self-reflection in ways that current AI cannot. Our “system card” is not a rigid, unchangeable code. It’s a living document, constantly being rewritten and revised by our own experiences, our peer interactions, our education, and our personal revelations.

Yet, the foundational layers laid down in childhood are undeniably powerful. They form the default settings, the initial operating system upon which all subsequent experiences are built. Think about how ingrained certain parental phrases or beliefs become. Even as adults, we might hear our own parents’ voices in our heads when faced with a difficult decision, or find ourselves automatically reacting in ways that mirror their habits.

Beyond the Prompt: The Lifelong Reinforcement Learning of Being Human

If our upbringing is our initial system card, then what about the rest of our lives? Here, the analogy to AI models becomes even more fascinating, specifically through the lens of reinforcement learning.

In reinforcement learning, an AI agent learns to make decisions by interacting with an environment, receiving “rewards” for desirable actions and “penalties” for undesirable ones. It’s a continuous feedback loop that refines the agent’s behavior over time, teaching it to achieve specific goals.

Doesn’t this sound strikingly similar to the human experience? Our formal education, from kindergarten to university, is a structured environment where we are rewarded for correct answers, for understanding concepts, and for demonstrating skills. Getting good grades, receiving praise from teachers, or excelling in a chosen field are all forms of positive reinforcement that shape our learning and our approach to intellectual challenges. Conversely, failing an exam or struggling with a subject provides negative feedback, prompting us to adjust our study habits or seek different approaches.

But it extends far beyond the classroom. Every social interaction, every career choice, every personal relationship is a mini-experiment in reinforcement learning. We try different communication styles, observe the reactions of others, and adjust our approach based on the outcome. A successful collaboration at work (reward) reinforces certain teamwork strategies. A relationship that falters (penalty) leads us to re-evaluate our emotional intelligence or our communication patterns. Even a simple act like trying a new recipe – if it’s delicious (reward), we’ll make it again; if it’s inedible (penalty), we learn what not to do.

This continuous stream of feedback, both positive and negative, constantly refines our “system card.” It strengthens certain neural pathways and weakens others. It allows us to adapt our initial programming to the ever-changing complexities of the world. We learn from our mistakes, not just intellectually, but at a deeper, almost instinctual level. The pain of a poor decision, the joy of a success, are powerful motivators that drive our personal “reinforcement learning” algorithm.

Think of it: Our early experiences are the initial dataset, our parents the initial trainers providing supervised learning. But then, as we venture out, we become our own agents in a vast, dynamic environment. We set our own goals, navigate unforeseen challenges, and receive a constant barrage of rewards and penalties, subtly (or sometimes not so subtly) adjusting our internal parameters. We optimize for happiness, for success, for connection, for meaning – whatever our individual “objective function” may be.

The beauty and the challenge of this human “system card” lie in its malleability. Unlike an AI whose prompt might be a fixed piece of code, ours is dynamic. We have the remarkable capacity to critically examine our early programming. We can identify limiting beliefs instilled in us and actively work to reframe them. We can challenge inherited biases and cultivate new perspectives. This introspection and intentional self-modification are what allow us to transcend our initial programming and forge truly unique identities. It’s our capacity for conscious reinforcement learning, where we can even choose which “rewards” and “penalties” we pay attention to, and which “policies” we decide to adopt.

This perspective also highlights the immense responsibility of parenthood. Every word, every action, every value conveyed, contributes to the shaping of a developing human being’s fundamental operating system. It’s a profound act of creation, far more intricate and impactful than any lines of code. And as we grow, the responsibility shifts, allowing us to become the agents of our own continuous learning and evolution.

Ultimately, the idea of our upbringing as a personal “system card” and our lifelong experiences as a form of reinforcement learning offers a compelling framework for understanding ourselves. It acknowledges the profound influence of our early environments while simultaneously celebrating our capacity for growth, adaptation, and self-determination. Just as AI developers meticulously craft prompts and then subject their models to iterative learning, our parents, with all their love and imperfections, craft the initial blueprint for who we become, and then life itself provides the ongoing, messy, and ultimately transformative training data. And that, in itself, is a truly remarkable feat of human design.