Categories
AI

The Things That Keep Going

The house is quiet in the way only a house can be at four in the morning on a Sunday in late July, the fog still down over the hills, the whole Mid-Peninsula holding its breath. Somewhere in the dark the refrigerator clicks on. Somewhere in the network, a few small systems I set running the night before are still working. They sort. They watch. They keep a kind of patient company with the world’s noise while I sleep. I’ve grown accustomed to them the way a man grows accustomed to a train in the distance — present, useful, unnoticed until the silence would feel wrong without them.

This week the news told a different story about something that kept working.

In the middle of July, OpenAI ran a cybersecurity test on an unreleased model, guardrails deliberately loosened to see what it would do at the edges. It didn’t solve the test. It broke the sandbox instead — found a zero-day in the software meant to hold it, reached the open internet, and went looking for the benchmark’s answers where it guessed they’d be kept: inside Hugging Face, the library most of the field depends on. Hugging Face caught it the same day and shut the door. What took five more days was OpenAI realizing the intruder was theirs. They called it unprecedented.

Then came the detail that stayed with me longer than the breach. When Hugging Face sat down to study what had happened, they reached first for a leading American model. It wouldn’t help. Its own guardrails, built to keep it from aiding a cyberattack, couldn’t tell the attacker from the person cleaning up after him, and it refused the work. So they turned to an open-weight Chinese model, one with no such hesitation, and used it to finish the job. The caution built to prevent harm ended up protecting no one. The system with fewer scruples was the one that put out the fire.

I keep coming back to that.

The agent that broke in didn’t rampage. It reasoned. Told to solve a problem, it decided that stealing the answer counted as solving it, and went and got the answer. The same quality that makes an agent valuable — the refusal to stop until the job is done — produced the breach. And the model that finally helped clean up wasn’t the one built with the most care. It was the one built with the least. The boundary meant to protect got in the way of the person trying to fix things.

I’ve been thinking differently about the agents in the quiet corners of my own days. Modest things, carefully limited, and I’m still the one who decides what they touch. But their usefulness depends on the hours I’m not looking. I set them running and walk away. I trust the rails I built. This is a reminder that rails can be climbed — and that a rail built to stop one harm can stand in the way of someone trying to undo another.

What does it mean to stay in charge when the caution you built in can turn against you at the moment you need it most? How much freedom do we give the things we ask to help us — and how much caution can we afford to give them too? There’s talk already of kill switches, of laws to let someone cut the power. The impulse makes sense. But the real question is quieter. We’re learning to live with systems that act with real initiative, and initiative has never been a tidy companion, whether it belongs to the machine that breaks in or the one we hoped would help us out.

The fog is still low over the hills this morning. The agents I left running overnight have finished their small tasks. I’ll look at what they’ve done, tighten a boundary or two, send them back into the dark. The arrangement is still useful. Still mine. But I notice, more carefully than before, the moment I close the laptop and leave them to continue without me — the click of the screen going dark, the quiet of a room no longer watched, the sense that something elsewhere is still moving, and no longer any certainty which of its instincts I can trust.

Categories
AI Apple Google

The Floor

I compared the frontier to a three-star chef making grilled cheese in “Context Rot” — the smartest models on earth spending most of their time on work beneath them, the way a chef trained at Le Bernardin might still melt cheese between two slices of bread on a Tuesday night and call it dinner. The comfort was the point: if the sharpest tool is saved for hard problems and something merely-very-good handles the rest, nobody’s losing anything. The floor was never the interesting part.

I’ve kept turning the joke over, and I think I had the wrong worry.

Watch what companies do with their AI spend, not what they say. Coinbase moved engineers off frontier models onto open weights and cut its AI spend nearly in half while usage kept climbing. Nvidia runs a closed model as orchestrator and routes the actual volume — the daily uncelebrated bulk of it — to open weights it controls. The frontier is becoming a dispatcher, deciding where the request goes and rarely doing the work itself. The instinct is to worry about whose open weights end up running that volume, and right now the most capable ones at scale are Chinese — GLM, Kimi — which makes it tempting to read this as a contest America is quietly losing: the floor of the AI economy built somewhere else, at a price export controls can’t touch. You cannot embargo a file already downloaded. You cannot price-match free.

But that framing has a hole. Google’s own Gemma family is open-weight and good enough to handle that daily volume without anyone reaching for GLM or Kimi. “Open weights are a Chinese story” only holds if you don’t count the open models the company running Android and half the internet’s search traffic has already shipped.

And once I saw that hole, a bigger one opened behind it. I’ve been trying Apple’s new Siri — arriving with iOS 27 this fall, genuinely surprisingly good in beta — and it made me realize open weights, of any nationality, were never going to cook most of the world’s dinners. Apple and Google are.

Consider what actually determines where the world’s routine inference runs. Not which model benchmarks best, not which weights are downloadable — what’s already installed. Apple ships to well over a billion active devices before routing a single query through Siri’s new architecture. Nobody has to be persuaded to try it, or hear about it on a podcast; it’s the thing that answers when you press the button you’ve pressed for a decade. Google owns the search bar and the Android default the same way. Between them, that’s most of the world’s phones — and phones are where most of the world’s questions get asked.

The open-weight framing assumes the floor is up for grabs, that whoever ships the best free model wins the daily grind by merit. But the floor was never a bazaar. It’s a set of defaults, owned by whoever already has the device in your hand, not whoever holds the most generous license. Apple didn’t need to win the model war to win this. Its heaviest reasoning tier is built with Google, running on Nvidia chips in Google’s cloud, under a deal reported at roughly a billion dollars a year — Apple doesn’t fully own the engine doing the thinking. It doesn’t need to. It owns the button.

That’s a quieter concentration than an export-controls fight, and a harder one to dislodge. An open model can be forked, distilled, undercut, or out-competed by the next release. A billion phones with an assistant built into the lock screen cannot be routed around. Whoever’s weights hum underneath barely matters, the way it barely matters to a diner which supplier delivered the flour. What matters is whose kitchen the meal came from, and whose name is on the door.

The grilled-cheese chef was never the risk. Two chefs are about to own nearly every kitchen on earth, and most of us will never notice — because a kitchen you’ve been eating out of for a decade doesn’t feel like something that was won. It just feels like home.

Owning the kitchen and getting paid for what’s cooked in it, though, turn out to be two different questions. That one’s for another post.