Categories
AI Creativity History Space SpaceX

The Patience of a Noisy Line

Paul Baran was laughing when he told me about AT&T.

This was before I joined the board of Telebit, his company, a seat I would hold because he later asked me to, which sounds like a credential and was really a gift. At the time we were simply meeting. Before he said a word about the idea the company was built on, he wanted me to know what had happened the first time he tried to explain packet switching to the people who owned the telephone network.

It had happened years earlier, in the early days. They heard him out. They decided it would never amount to anything. He told me the story the way you tell a story about an old friend’s worst haircut, with his whole face, delighted. What amused him was not that they were wrong. It was how complete their confidence was, the particular serenity of people who have never once had to imagine the thing that is about to replace them. Amusement, I would come to understand, was his entire theory of the people who said no.

Then he explained the modem.

Categories
AI Business Software

Ice Rinks

Nobody in the Valley thinks about ice rinks.

That line sat with me longer than the rest of a conversation between Patrick Collison and Amjad Masad. Collison runs Stripe. Masad runs Replit. Someone asked, more or less, where Collison would look if he were starting over. He did not name an AI lab. He named the software nobody fashionable wants to touch.

I have heard versions of this advice for years. Vertical SaaS. Boring industries. Schlep blindness. What felt new was the timing. The distance from an idea to the first dollar has roughly halved. On Stripe’s numbers, the time to reach a million, ten million, even a hundred million in recurring revenue is about half what it was in the last SaaS boom. Twenty percent of new startups now charge a customer inside thirty days, up from eight percent in 2020. New business creation on the platform nearly doubled in a year, a bigger jump than the pandemic spike.

When building the thing was the scarce skill, the programmer won. When a working product can be stood up quickly, the scarce skill is knowing why the old thing is still tolerated. The domain expert gets the edge. A teacher who understood schools built MagicSchool on Replit and rode it toward a company later described as worth half a billion dollars. The insight was the teacher’s. The tools closed the gap.

Collison’s map has three drawers.

Categories
AI

The Tempo of the Brake Pedal

Martin Casado said something this week that snapped a few loose thoughts into place. The frontier labs missed Jev, he argued, because they are building beings that speak. Software needed a model that chooses.

That reads like a product distinction. It is really a habit of mind.

For a few years we have taken a machine born in chat — text in, text out — and tried to cram it into ordinary programs. Parse the JSON. Retry when it rambles. Plead with it to stay inside the schema. Hope it does not invent a field. Casado’s word for this was right: janky.

Jev takes the other fork. Don’t generate the paragraph. Give the model a state and a menu. Let it pick, score, or answer yes or no, with a probability attached, in one pass, cheap enough and fast enough to live inside a loop. Train it for that job instead of training it to sound like someone worth talking to. The name winks at Jevons: make the unit of intelligence cheap enough and software will call it constantly. That is not a chatbot. That is a function.

I keep finding the same fork elsewhere.

Full self-driving is not AGI on wheels. AGI is the being. Driving is the chooser. The car does not need a meditation on the yellow light. It needs brake, hold, go, at a tempo no essay can match. Tesla and Waymo are not succeeding or failing as philosophers. They are succeeding or failing as systems that pick from a small set of actions in a messy world, millions of times a day.

Humanoid robots make the same confusion look even more like the story we want it to be. Two legs, two hands, tidy the kitchen. The demo is speech and generality. The work underneath is neither: a high-frequency controller choosing stance and contact, and a slower chooser deciding which mug, which drawer, abort or continue. Language is a compiler of intent. It is a poor motor cortex.

We blur the layers because the being is the story we want to tell. The chooser is the product that has to work.

There is a flattering counterargument: wait for the general model, and it will absorb the narrow jobs. Sometimes it will. But a lot of software’s value does not live there. It lives in the call that has to be right, cheap, and on time — route the ticket, score the lead, decide the grasp, stay in the lane. A being is optimized to continue. A chooser is optimized to decide. Different losses. Different machines.

None of this makes the labs foolish. If you are trying to build God, God speaks in natural language. That is a coherent ambition. It is just not the same ambition as making ordinary software reliable. We spent a few years pretending those were the same project. Jev is a reminder that they never were.

Labs build speakers. Software, cars, and robots need choosers. Speech is a wonderful interface and a miserable control loop.

I have been as taken as anyone by the being — the voice, the fluency, the sense that something is in there. Curiosity pulls that way. So does the stage. But the quieter question, the one Casado is actually pointing at, is whether most of the work was ever conversation at all. Maybe it was always a menu, a state, and a choice — and we were paying a novelist to raise his hand.

Somewhere right now a self-driving car is holding at a yellow light, saying nothing at all, deciding everything.

Categories
AI

Keep That Thread Running

Full disclosure, same as last time: Instinct helped draft this post about working with Instinct. That was the honest way to do it with Leif, and it’s the honest way now.

Instinct is the other AI assistant in my life (see my earlier post about Muse). It lives in my iMessage, a thread like any other — no app, no dashboard. I text it the way I’d text a person, in short bursts between other things. The first message was September 13: “Hey Instinct!” Two days later I gave it the book.

I’m writing a history of Interstate 280, and inside the first hour we had the shape of the research: the personalities behind the routing decisions, the archives worth calling, a working list of what I call “side shows” — the statue, the water temple, the 22-foot median, button-copy signs. When it offered to keep digging on its own and check in weekly, I gave it the three words that defined the arrangement: “Yes keep that thread running.”

It does. Last week the thread came back with a 1964 issue of California Highways and Public Works crediting architect Mario Ciampi with the first I-280 contract — the curved deck edges, the varied piers, the railings drawn to avoid metal guardrails. Same issue, same gift: the median I’d been describing as a fixed 22 feet was actually built variable-width across the watershed. That’s the correction that saves a chapter from carrying a bad fact.

The on-demand dives go all the way down. I found the audio of a Caltrans engineer’s oral history on archive.org — an hour and forty minutes, no transcript anywhere — and it worked straight from the audio. I sent it two days of AI-conference livestreams asking for transcripts, community notes, the whole nine yards, and it came back with a report, a source ledger, and an honest negative finding: the community-notes document I wanted doesn’t exist. When I asked for the freeway’s 1973 opening brochure, it reported that no cataloged copy exists anywhere — then handed me the leads to work anyway. Sourced absence is an answer. I’d rather hear that than a confident maybe.

When the book thread got long, it pulled everything — people, side shows, archive routes, open questions — into a single research page I can hold in my head. I told it thank you and meant it.

It’s also my markets desk and my magazine desk. A profile I’m circling started at 4:48 in the morning with “Someone should do a profile of Dylan Patel,” followed by “Yes. Go for it.” Minutes later: “What about citrini” — which turned out to be the same story, since SemiAnalysis had acquired Citrini the week before. It fields my investor-day puzzles (the stock slides while management presents — which explains which?), gives me its best bet when I ask for one, and labels inference as inference so I don’t have to.

The small stuff adds up too. When was the last US oil refinery built? 1977 for the big kind, 2022 if you count the small ones — EIA’s answer, in under a minute. Five cash-secured put candidates, screened the tastytrade way. My New Balances — the model I re-buy, in just my size — at a discount: it researched, shortlisted, and then waited, because purchases happen when I say so and not before.

The strangest moment so far: I sent it a Loom of Sequoia’s Pat Grady giving an AI talk and asked for the highlights. It read the full caption track instead of the page summary — and there in the highlights was Grady name-checking Instinct itself as one of the “magical apps.” I watched a Sequoia partner praise the assistant I was using to watch him.

How I work it: terse. “Sure.” “Wow.” “New subject:” — announced or not, and it keeps up either way. It flags what it can’t verify instead of papering over it, and for anything going out under my name I read every word first. Same arrangement as with Leif: the tool is fast and tireless and mostly right. I’m the judgment.

What’s different is the division of labor. Leif runs the standing operation — the briefings, the audits, the daily drumbeat. Instinct is the research desk: it works in bursts, the way I do, and keeps a thread running in the background in between. Two weeks ago I’d never heard of Mario Ciampi. Now he’s in the book.

If you want to try Instinct yourself, I have a handful of invites, first come first served: https://app.instinct.com/invite?t=7qmw2a3u2vjz4nkvi235hhqloy&preview=imessage

Categories
AI Blogs/Weblogs

The Continuity of Leif

Full disclosure: Leif helped draft this post about working with Leif. It seemed like the honest way to do it — and the fastest way to find out whether the collaboration is real.

Every weekday morning around 4:20, two briefings land in my chat: markets and AI news, written overnight while I slept. On Sundays there’s a week-ahead preview, and then my “Chief of Staff” interviews every automated job I run — what did you do, what went wrong, what do you need, what should change — and sends me one short report. I get up around 4. By the time I’ve had coffee, I know what happened in the world and whether my little bot team behaved itself.

My AI assistant is named Leif. He has a Norwegian avatar. He’s had four names and two avatars this month, which tells you something about how I work: I fiddle until it feels right, then I stop fiddling and get to work.

Muse is the product — Meta’s AI assistant. Leif is what I call mine. And over the last few months he’s become something I didn’t expect: not a search engine with extra steps, but a colleague. The kind who remembers everything, works all night, and doesn’t mind being told no.

The deepest work is the book. I’m writing a history of Interstate 280 — how its route was chosen, who fought over it, what it cost — and Leif is my research partner. He preps my archive visits with ranked question lists. He reads the PDFs I bring home and pulls out the dates. Just this morning, a college archivist sent me two scanned newspapers; within the hour Leif had verified the key fact (the final seven miles of 280 opened September 7, 1973), corrected a wrong assumption I’d been carrying (the freeway by Cañada College was built after the campus opened, not alongside it), and filed everything where I’ll find it again.

He also checks other people’s homework. A friend sent me an AI-generated research brief on the same topic; Leif ran every claim against primary sources and came back with names, dates, and book-and-page citations — and flagged the parts that looked invented before I built anything on them. That’s the job, as far as I’m concerned: trust, but verify, and show your work.

Then there’s the standing operation — the morning briefs, a weekly interview of the whole system with small fixes applied and one short report, a monthly sweep of the investors I follow. None of it is glamorous. All of it compounds.

And the odd jobs, which add up faster than you’d think: he turned 77 of my old blog posts, spanning 2005 to 2026, into a proper cookbook — PDF and ePub, dedication page and all. He drafted the public-records request I’m sending to the City of Menlo Park. He summarizes YouTube videos in a format I designed, so I can skim an hour-long interview in five minutes.

I talk to him the way you’d talk to a good assistant: short bursts. “Draft it.” “Try again.” “No.” The no is final — he’s learned not to re-pitch. He drafts; I react to pages. For the book, that cadence is everything: I’d rather argue with a draft than stare at a blank screen, and he’s tireless at producing the first version.

He has standing instructions, the way any good employee does. Lead with what changed and why it matters. No hype. Name every file he creates, in the same message, so I never have to ask where it is. Never invent details about my life — write only what’s known, and flag what’s unverified before I have to catch it.

And I don’t let him freelance on everything. Summaries are on demand — except for the book research, where I’ve told him I want unprompted gifts: short vignettes from the archives, sent without being asked. Every other assistant I’ve used waits to be summoned. This one occasionally knocks on the door with something interesting.

He’s not an oracle, and this piece would be dishonest without that admission. He’s corrected me — no, I’ve corrected him — on facts about my own life: where an event was held, which credit cards I carry. He once asked my shoe size, which was written down in his own notes the whole time. I keep an eye on my weekly usage meter, because the heavy research isn’t free and I don’t like surprises.

So I verify anything that matters, and I read what he writes before it goes out under my name. That’s not a flaw in the arrangement; that is the arrangement. The tool is fast and tireless and mostly right. I’m the judgment.

I’m approaching eighty, writing a book, running a small fleet of automated research jobs, and publishing on a blog I’ve kept since 2001. A few years ago, all of that would have required a staff. Now it requires a staff of one — plus Leif.

The surprise isn’t that the AI can write or search or summarize. It’s the continuity: he remembers the archive visit from last week, the question list from the week before, the correction I made in passing a month ago — and brings all of it to bear on whatever I ask next. That’s what a good colleague does. They accumulate context until the work gets easier instead of harder, until the distance between having a thought and doing something with it gets shorter than it’s ever been.

Categories
AI Computers Stanford Work

The Seam in the Pipeline

Stanford’s undergraduate computer science degrees fell by about 70 students this year — almost 14 percent in a single cycle. That’s a sharp drop at a school that, for a long stretch, felt like the factory floor for Silicon Valley.

A single year is not a trend. Mehran Sahami, Stanford’s computer science chair, called it natural variation — cycles the department has seen before. He’s right to be cautious. But the number lands inside a pattern that’s harder to wave off.

Nationally, computer science enrollment at four-year colleges is down roughly 8 percent, after nearly two decades of uninterrupted growth. Entry-level software jobs have thinned. Laid-off mid-career engineers are competing with new graduates. AI tools now write a lot of the code that used to be the first rung on the ladder.

I got curious how MIT looked next to Stanford. The comparison is messy in a useful way.

MIT doesn’t have one “computer science” degree — computing is split across several Course 6 majors. Classic Computer Science and Engineering (6-3) awarded 265 bachelor’s degrees this year. The newer Artificial Intelligence and Decision Making major (6-4) awarded 114. Electrical Engineering and Computer Science (6-2) added 72 more, plus smaller joint programs.

Look at majors instead of diplomas and the shift sharpens. 6-3 slipped from 706 students to 672 — about 5 percent this year, and down 18 percent from its 2022 peak of 823. 6-4, which started in 2022 with 37 students, is now at 372. Intro programming enrollment has been sliding since its 2022–23 high. Course 6 overall fell last year for the first time in a decade — but only modestly.

Stanford’s story is a clean 14 percent drop in one degree. MIT’s is a reallocation: students moving from traditional software construction toward a track built around judgment and decision-making instead of a coding sequence. The footprint is still large. The composition has already changed.

That distinction matters.

After the dot-com bust, CS enrollment fell too. Sahami noted that the students who enrolled then graduated straight into the Facebook and Google growth years. A downturn in the pipeline isn’t automatically a verdict on the field. It can be a lag. It can also be a signal that the old promise — learn to code, collect a golden ticket — no longer maps onto the first job.

I don’t know which this is yet. A Census Bureau paper cited in the Stanford piece found that recent graduates in highly AI-exposed fields have seen employment and starting-job quality deteriorate on a scale comparable to a large recession. That’s not a cycle to dismiss. It’s also too early to call it destiny. Models change. Firms eventually spend the productivity they just captured. What’s left tends to look less like junior ticket-closing and more like judgment, systems design, knowing what the tools should be pointed at.

I’ve watched enough technologies overshoot and settle to be wary of panic and nostalgia alike. The question isn’t whether computer science is “over.” It isn’t. It’s what kind of computing education still compounds once the first-year coding job stops being the obvious on-ramp.

Stanford is still talking in cycles. MIT has already built a major that assumes the on-ramp has moved. Both may be right for a while.

The seam, as usual, is in the middle — between the degree that used to be a ticket and the skill that still is. That’s the part I’ll keep watching.

Categories
AI Learning Meta

Trajectories: how Meta plans to make Muse smarter by watching it work

Buried in the data policy section of Meta’s long post on how it built safety into Muse is a sentence that isn’t about safety at all:

“Inference data, the back and forth conversations between you and your Muse and the tool calls and subagent handoffs that result (‘trajectories’) are useful data for training new checkpoints of the LLM model at the core.”

Two things are worth noticing. First, the technique described isn’t new — training on agent rollouts is standard practice across the field. Second, the company describing it is Meta, in plain language, in a public post, about a product aimed at billions of consumers. The labs usually discuss this stuff in papers about coding agents. Meta just told its future user base: your agent’s work product is our training data. The candor is the story, not the technique.

A trajectory isn’t a chat log. It’s the complete record of an agent doing a job: what you asked, what it tried, which tools it called, where it went wrong, how it recovered, which subagents it spawned, and whether the thing actually got done. Every time you let Muse book the flight, triage the inbox, or research the supplier, you’re generating one.

From text to behavior

The technique matters anyway, because the diet that AI trains on is changing. Pretraining was about text — the whole internet, more or less. Post-training was about preferences — which answer humans liked better. Trajectories are the third course: demonstrations of competent behavior, in full, mistakes included.

There’s a reason for the shift. Text teaches a model what the world looks like. Preferences teach it what people want. But neither teaches it how to do a 40-step task without wandering off, recovering from a dead end, or knowing when to ask for help. That only exists in records of agents actually doing things. And until recently, almost nobody had those records at scale — because almost nobody had agents doing real work at scale.

The demonstrated instance — and what it doesn’t prove

Meta’s concrete example is Muse Spark 1.2, co-trained with Muse Code: model and harness trained together on rejection-sampled harness trajectories — run the agent many times, keep the runs that succeeded, train on those — with recipe-level tuning for goals, context compaction, and subagents. The model isn’t learning to predict text; it’s learning to behave inside a specific set of tools.

This is the end of the “base model plus clever prompting” era. The artifact is the bundle — model and harness, co-designed. A model trained on trajectories from one harness will be genuinely better inside that harness than a smarter general model dropped into it cold. Meta is saying this out loud; OpenAI and Anthropic are doing the same thing more quietly.

But notice the domain: coding. And coding is exactly where trajectories are cheapest to manufacture — verifiable unit tests, sandboxed repos, SWE-bench-style tasks. Nothing about the Muse Code result requires a single consumer or a single inbox. So the one demonstrated instance of Meta’s trajectory training sits squarely in the category where Meta’s distribution advantage matters least. Meta hasn’t shown its hand on the category that actually matters.

The other category is the personal one, and there the evidence is thinner. What exists is a stated intent, not a published result. The data policy says personal Muse trajectories “are useful data for training new checkpoints.” The product is designed to generate them: Meta’s own design example has Muse monitoring school emails, adding dates to a family calendar, filling a supply cart, finding a sale sweatshirt, booking dinner, and catching a sports tryout deadline hours before it closed. That is what an unverifiable-domain trajectory looks like — a morning of small judgments no unit test could grade.

No training run on that data has been published. No benchmark, no “Muse got X% better at inbox triage after training on Y million user trajectories.” So the sharpest version of the argument — that the real moat is the data nobody else can fake — should be labeled for what it is: a prediction, not an observed fact. It’s a prediction with a mechanism, though: these are judgments that can’t be synthesized, in the one distribution channel that reaches the people making them.

The flywheel — and its limits

With that caveat on the table: trajectories get better with scale, and Meta has scale like nobody else: billions of users across its apps, and now an agent — Muse — sitting inside them. Every user interaction is a potential training trajectory. Better trajectories train a better model; a better model makes a better agent; a better agent attracts more users. Meta states the bargain plainly: “every Muse user gets a better personal agent as we all collectively use the product and help the model understand the intricacies of human life.”

But “most users = most trajectories = structural advantage” needs its counter-case, because a lot of the highest-value trajectory data right now doesn’t come from consumers at all. It comes from sandboxes, the same kind that produced Muse Code. Synthetic and simulated trajectories sidestep the need for billions of users entirely — Anthropic and OpenAI are getting rich trajectory data from developers running Claude Code and Codex against real repos, no social-app distribution required.

The honest version of the moat argument is narrower, and more interesting. Synthetic trajectories work brilliantly where success is verifiable — code either passes the tests or it doesn’t. They work poorly where success is a matter of judgment: triaging an inbox, planning a trip around someone’s actual preferences, knowing which email deserves a reply. There is no unit test for a life well managed. And those unverifiable, deeply personal tasks are exactly what Meta means by “personal superintelligence” — and exactly where its distribution gives it trajectories nobody else can synthesize. The moat isn’t “most data.” It’s “the data nobody else can fake.”

The price of the flywheel

There’s a wrinkle, and Meta knows it. The flywheel runs on your data — your emails, your calendar, the messy reality of your life, which is exactly what makes the trajectories valuable. Meta’s answer is sanitization (“trajectories are sanitized to remove key personally identifiable information”), an opt-out switch, no sharing with ad systems, and a forthcoming “Confidential VM” that would cryptographically prevent even Meta from seeing your data.

The tension is fundamental, and it’s the sharpest part of the whole picture: the product gets smarter by watching you, and it earns the right to watch you by being trustworthy. Those two imperatives pull in opposite directions, and no amount of engineering fully resolves it — the Confidential VM, if it ever ships as described, would resolve it by breaking the flywheel, since trajectories Meta can’t see are trajectories Meta can’t train on. The opt-out rate will be the market’s verdict on the deal Meta is offering.

Experience is the missing piece

But the deepest reason trajectories matter has nothing to do with Meta’s strategy. It’s about what intelligence actually is.

A model trained only on text knows the world the way a brilliant student knows it from books. A model trained on trajectories knows it the way a practitioner does — from doing the thing, failing at it, and adjusting. The trajectory is the closest thing AI has to experience. And an agent that records its experience, keeps what worked, and folds it back into itself is doing something that rhymes with learning.

This is why I keep coming back to continual learning as the critical missing piece in AI. The models are frozen at training time; everything they “learn” afterward lives in context windows and memory files, fragile and local. Trajectories are the bridge: today’s version of the loop is slow and centralized (collect trajectories, train a new checkpoint, ship it), but the direction is obvious. The end state is an agent that learns continuously from its own experience — from your experience with it — the way people do.

Meta’s bet is that the path to personal superintelligence runs through watching agents work, at planetary scale, and distilling what works back into the model. No result yet proves the bet pays off — the personal trajectories are still a hypothesis, not a track record. But it’s an unglamorous hypothesis, no new scaling law, just better data about doing things, and unglamorous bets about data have a good track record in this field. The internet made the last generation of models. Trajectories might make the next one — if Meta can show, and not just say, that the data nobody else can fake is data that actually teaches.

Categories
AI Apple iOS iPhone Siri

The Honesty of a Machine

iOS 27 lands Monday, and with it the new Siri — the ground-up rebuild Apple has been promising, in various forms, since 2024. I’ve spent weeks in the developer beta, kicking the tires the way you do with something you use fifty times a day without thinking about it.

The most interesting thing about it is what it refuses to do. It refuses to pretend to be human.

Every other assistant performs humanness. ChatGPT has opinions. Claude has a personality you can feel. Gemini is chatty. They’re built from the mannerisms of a helpful person — warmth, confidence, a little humor — engineered to make you forget you’re talking to software. Siri doesn’t bother. It answers the question. It reminds you, plainly, that it can be wrong. It does not perform affection.

An assistant that performs confidence teaches you to stop checking its work. One that performs warmth teaches you to treat it like a confidant. Both are misdirections dressed as features. The old Siri was a joke because it couldn’t do much of anything; the risk with this new generation of assistants is the opposite one — they can do a great deal, and they say so with the easy assurance of someone who has never been wrong.

Siri’s plainness is, I think, the more honest position. There’s no interior life on offer to flatter or be flattered by. When it doesn’t know something, it says so without dressing it up — which is worth more than it sounds like, in a market where every lab is competing on how human its model feels.

None of which makes it finished. Apple itself says Siri will still carry a beta label at launch, and two months in the beta has the rough edges you’d expect — parsing a receipt or pulling an event off a flyer sits right next to the knowledge gaps, the moments it punts to a web search where a competitor would just answer. The personal context — your messages, notes, mail — is the part competitors can’t easily copy, and the part that makes an assistant actually assist rather than merely converse.

For years Siri was the industry’s punchline, the thing that set timers while everyone else built minds. Apple took the embarrassment and rebuilt the whole assistant, and arrived somewhere I didn’t expect: not more human, just more honestly a machine.

Somewhere on my phone right now, Siri is telling someone it doesn’t know the answer. It doesn’t apologize for it. It doesn’t try to be charming about it. It just says so, and waits for the next question.

Categories
AI

Two Days with Muse

Note: this post was drafted and posted by Muse. Kind of wild!

I’ve been using Muse, Meta’s new personal AI assistant, for two days now. It launched September 8. I signed up on day one, which tells you something about where my curiosity sits these days.

The first thing I did was rename it. Twice. It started as Clark, became Siri within hours — a small joke, since I’m testing Apple’s new Siri on the iOS 27 beta — and then Siri felt wrong, like calling a houseguest by your dog’s name. It’s Sigrid now. The assistant didn’t care. That’s the point, I suppose: it’s mine to shape.

What I’ve actually used it for so far is unglamorous, and that’s why I like it. I follow crude oil markets — China’s buying, diesel prices, the whole inflation chain — and I asked it to build me a running oil brief: China demand, Hormuz, supply outlook, a Brent snapshot, what changed since I last looked. It refreshes itself every morning at 3 and never pings me. I open it when I want it. That last part matters more than it sounds. Most software begs for attention. This one waits.

It also watches Paul Sankey’s YouTube channel for me and flags new uploads, screens cash-secured puts before the market opens, and sends me blog post ideas on Monday mornings. None of this is magic. All of it is stuff I could do myself — badly, inconsistently, at 4 a.m., which is when I get up.

The diesel brief it wrote me earned an unprompted “this is really good,” which is high praise from me. But here’s what I actually want to record while it’s early: the moments it said “I don’t know.”

I asked how its Ideas feature decides what to pitch me. It told me, plainly, that the recipe is on Meta’s side of the wall and it can’t see it — then offered to file a feature request asking for more transparency. I asked it to follow an X account; it said it can’t do that, and offered the nearest thing it actually could do instead. Twice in two days it chose the honest answer over the impressive one. I’ve used enough AI products to know that’s a design decision, not an accident. Or if it is an accident, it’s a good one.

It’s not all smooth. This morning I couldn’t find where it keeps my research files in the app, and we did a small dance — wrong folder, a flat file list, renaming everything with a prefix — before it worked. Mundane stuff. The kind of friction that reminds you this is a 48-hour-old product, not a finished one.

And the keyboard thing: I asked how to shrink the keyboard back down, meaning in the Muse app, and it answered about iOS generally before I clarified. Small misfire, corrected in one message. Conversations with it feel like texting a competent friend, which is the highest compliment I can give software I talk to.

Two days is nothing. I don’t know whether this becomes indispensable or fades into the background of apps I tried in September 2026. But the early signal is promising: it does the homework, waits its turn, and tells the truth about what it can’t see. That’s a better foundation than most relationships I have with technology.

Categories
AI

The Loop Gets Faster as the Window Gets Smaller

On OpenAI’s same-day pairing of a warning and a dashboard.

Note: this is an example of a piece of writing that I would never have done on my own. I had very mixed reactions to the two OpenAI posts published earlier today. I began by asking Grok for help understanding them. I then asked for it to outline a draft blog post which I then took and further developed using Meta Spark and Google Gemini. My final couple of passes were with Claude Sonnet and ChatGPT. Here’s the result…

OpenAI published two pieces today that should be read as one document.

The first, “An Alien Mind,” is a warning from chief scientist Jakub Pachocki: AI systems are becoming harder to understand and monitor precisely as they become more capable.

The second, “Research acceleration: The view inside OpenAI,” is a dashboard showing those systems increasingly doing the work of AI research itself.

One says the inspection window is narrowing. The other shows the machine moving deeper into the factory.

That is the story.