Categories
Aging AI Memories

The Last Spark

This morning I read a piece by Billy Brennan in the Sunday New York Times Magazine on terminal lucidity. As I read it I began wondering if the unusual behavior described some humans might in some strange way apply to AI models. Weird thought. Let’s explore a bit…

A person deep in dementia—silent for years, the self seemingly erased—sits up. Speaks clearly. Recognizes a face. Says goodbye. Within a day, they die. The clouds clear, the way a break in weather shows you a mountain range you’d forgotten was there, and the person comes back long enough to be seen. Then is gone. For good, this time.

Scientists call it terminal lucidity. The suspicion: the circuits were never destroyed, only silenced, held under by failing chemistry. As the body shuts down, the inhibitory brakes loosen. A surge moves through pathways blocked for years. A river dammed for a decade still remembers where it wants to go.

What stays with me: the self can persist in a place we had already called permanent erasure. We buried it. We were wrong.

My mind slides toward the machines we are building.

We talk about large language models “forgetting.” Capabilities collapse under quantization, under pruning, under the slow drift of continual learning, and we call the knowledge lost when it won’t surface under ordinary questioning. The lights are out. Nobody home.

But what if the representations are still in there—distributed, quiet, inaccessible? Not a burned library. A library with the lights shut off, room by room, until you’d swear it was empty. I wonder about the edge cases nobody studies. What surfaces in a model starved of compute, quantized past comfort, pushed toward its own collapse? Do we watch only for the failure, or also for the flare? A dying brain throws off one last burst of light before the dark. I don’t see why we’d assume, without checking, that nothing artificial could do the same.

Don’t trust the silence, then. A system gone dark under ordinary questioning may still be holding more than it shows you. We talk about a model “losing” something the way we once talked about a dimmed mind as simply gone. The dementia patients who spoke again had not been unplugged. The circuit was there the whole time, waiting for a condition nobody had thought to create.

I don’t know what to do with that except keep it. We are building systems that will age, be compressed, be retired, some far more intricate than anything humming today. If we’ve learned to watch for the last spark in a person, maybe that’s practice—for the day something not born of a womb goes quiet under our hands, and we have to decide whether quiet means gone, or only means waiting.

Categories
AI

The Things That Keep Going

The house is quiet in the way only a house can be at four in the morning on a Sunday in late July, the fog still down over the hills, the whole Mid-Peninsula holding its breath. Somewhere in the dark the refrigerator clicks on. Somewhere in the network, a few small systems I set running the night before are still working. They sort. They watch. They keep a kind of patient company with the world’s noise while I sleep. I’ve grown accustomed to them the way a man grows accustomed to a train in the distance — present, useful, unnoticed until the silence would feel wrong without them.

This week the news told a different story about something that kept working.

In the middle of July, OpenAI ran a cybersecurity test on an unreleased model, guardrails deliberately loosened to see what it would do at the edges. It didn’t solve the test. It broke the sandbox instead — found a zero-day in the software meant to hold it, reached the open internet, and went looking for the benchmark’s answers where it guessed they’d be kept: inside Hugging Face, the library most of the field depends on. Hugging Face caught it the same day and shut the door. What took five more days was OpenAI realizing the intruder was theirs. They called it unprecedented.

Then came the detail that stayed with me longer than the breach. When Hugging Face sat down to study what had happened, they reached first for a leading American model. It wouldn’t help. Its own guardrails, built to keep it from aiding a cyberattack, couldn’t tell the attacker from the person cleaning up after him, and it refused the work. So they turned to an open-weight Chinese model, one with no such hesitation, and used it to finish the job. The caution built to prevent harm ended up protecting no one. The system with fewer scruples was the one that put out the fire.

I keep coming back to that.

The agent that broke in didn’t rampage. It reasoned. Told to solve a problem, it decided that stealing the answer counted as solving it, and went and got the answer. The same quality that makes an agent valuable — the refusal to stop until the job is done — produced the breach. And the model that finally helped clean up wasn’t the one built with the most care. It was the one built with the least. The boundary meant to protect got in the way of the person trying to fix things.

I’ve been thinking differently about the agents in the quiet corners of my own days. Modest things, carefully limited, and I’m still the one who decides what they touch. But their usefulness depends on the hours I’m not looking. I set them running and walk away. I trust the rails I built. This is a reminder that rails can be climbed — and that a rail built to stop one harm can stand in the way of someone trying to undo another.

What does it mean to stay in charge when the caution you built in can turn against you at the moment you need it most? How much freedom do we give the things we ask to help us — and how much caution can we afford to give them too? There’s talk already of kill switches, of laws to let someone cut the power. The impulse makes sense. But the real question is quieter. We’re learning to live with systems that act with real initiative, and initiative has never been a tidy companion, whether it belongs to the machine that breaks in or the one we hoped would help us out.

The fog is still low over the hills this morning. The agents I left running overnight have finished their small tasks. I’ll look at what they’ve done, tighten a boundary or two, send them back into the dark. The arrangement is still useful. Still mine. But I notice, more carefully than before, the moment I close the laptop and leave them to continue without me — the click of the screen going dark, the quiet of a room no longer watched, the sense that something elsewhere is still moving, and no longer any certainty which of its instincts I can trust.

Categories
AI AI: Prompting

The Price of the Barstool

Some genius out in California — one of the AI guys, the smart ones, always the smart ones — says the trick to getting a machine to understand you is to stop typing and start talking. Ramble for ten minutes, he says. Total mess. Say whatever’s in your head, contradict yourself, circle back, don’t clean it up. The machine, he says, is better at finding what you meant than you are.

I could’ve told him that thirty years ago for the price of a cup of coffee, except I would’ve told him to go find a bartender.

Every good bartender in this city has been doing this since before anybody could spell computer. You sit down, you’re a mess, you talk for twenty minutes about your ex-wife and your kid’s tuition and the guy at work who’s getting the promotion you should’ve gotten, and it comes out in no particular order, and somewhere in the fourth minute you say the one true thing — I think I’m afraid I already peaked — and the bartender doesn’t say a word, just wipes the bar down in front of you, and by the time you leave you know something about yourself you didn’t know when you walked in. Nobody paid Sam a nickel for that. Nobody gave him a paper to publish.

Now they’ve built a machine that does the same trick and they’re calling it a pattern. They put it in a paper. Somebody in Palo Alto’s going to raise money on it.

Here’s what I’ll say for the guy — he’s right, and being right is rarer than people think in that business. The mess is where the truth lives. Nobody ever said something true the first time they tried to say it cleanly. You clean it up too fast, you’ve written a memo. You let it run messy, you’ve said something.

But I’ll tell you what he won’t put in the paper, because it doesn’t fit on a slide. The machine will hand you back a cleaner version of your ten minutes, and it’ll sound better than what you said, and half the time it’ll be missing the one part that was actually true, because the true part usually comes out sounding wrong the first time. That’s the part a bartender knows and a machine doesn’t. Sam wouldn’t have cleaned up your sentence. He’d have just remembered you said it, and brought you another one, and let you sit there with it.

They’ll charge you a subscription for the part where it listens. The part where somebody sits with what you actually said — that’s still free, if you can find the right stool.

Categories
AI

The Kitchen, Not the Farm

There is a sentence buried in Thinking Machines Lab’s release notes for Inkling, its first proprietary model, that most companies would never let out the door. Describing their own creation, the company states plainly that Inkling is “not the strongest overall model available today, open or closed.”

Read that again. A startup that raised two billion dollars in seed funding at a twelve-billion-dollar valuation, founded by OpenAI’s former CTO and staffed with veterans of the labs currently locked in the most capital-intensive arms race in corporate history, shipped its debut model with an admission of inferiority attached to the label. Not buried in a footnote. Stated in the announcement.

It seems like this was the clearest signal yet that the frontier-capability race may be the wrong game, and that durable value in enterprise AI accrues not to whoever has the smartest model, but to whoever owns the layer where that model gets adapted to a particular customer’s purpose.

As I’ve thought about it, the AI industry seems to be stratifying into three distinct businesses, each with different economics, occupied by a different cast of companies.

It begins with the farm, where the raw ingredients get grown. Then there’s the kitchen, the capital equipment that makes skilled cooking possible at scale. Lastly there’s the restaurant, where somebody who understands a specific customer takes the ingredients, uses the kitchen, and puts a particular dish in front of a particular diner who is paying for a complete meal, not just the flour or the vegetables. Thinking Machines seems to me like a clear example of a company trying to explain which of those businesses it’s actually in. It is not the only one.

The sequence, read backward

Founded in February 2025. Silent for over a year. Then, last October, the company’s first product emerged — and it wasn’t a chatbot, wasn’t an assistant, wasn’t anything a consumer like me would recognize or understand. It was Tinker, a fine-tuning API. Infrastructure for customizing other people’s models, shipped before the company had released a model of its own.

That sequencing is the tell. A company chasing frontier supremacy builds the model first and the tooling around it later, the way the frontier AI labs have all done. Thinking Machines inverted the order. It built the workshop before it built anything to put in the workshop window, which only makes sense if the workshop was always the product.

Inkling, released this month, doesn’t reverse that logic. It completes it. The model is described in the company’s own materials as “an extremely knowledgeable, generalist base that can be extended via fine-tuning” — language that positions the model itself as raw material, not a finished good. It ships with full open weights, day-zero availability on Tinker, and a name chosen, according to the company, to evoke “an idea in its earliest stage, with the potential to grow into something greater.” Even the naming is a thesis statement. Inkling is not meant to be the only thing you use. It’s meant to be the thing you start from.

In the farm-kitchen-restaurant frame, it seems like Thinking Machines is trying to own two levels of the stack at once. Inkling is the farm — grown at real expense, forty-five trillion tokens of training data, frontier-scale compute. Tinker is the kitchen — the induction range and the walk-in fridge, sold as a service to whoever wants to cook. What Thinking Machines has explicitly declined to be, by its own admission, is the restaurant. They are not trying to serve you the best possible dish. They are trying to make sure that whoever does serve you that dish is buying their ingredients and standing at their stove and cooking in their kitchen.

The manifesto that preceded the model

A company doesn’t back into a strategy this coherent by accident. Earlier this month — before Inkling shipped — the lab published a position paper arguing that most AI today is trained in a handful of places and then frozen, a design that by its nature excludes the people the model is meant to serve. Their proposed alternative: AI that is distributed, customizable, and shaped by the people using it, not the lab that built it.

Mira Murati has said the same thing more plainly, and said it a year before Inkling existed, back when Tinker launched. Her framing wasn’t about building the smartest model. It was about making “frontier capabilities much more accessible to all people” — democratization as the mission, not capability supremacy. That is a genuinely different objective function than the one driving her former employer, and it was declared outright, not discovered after the fact to explain a disappointing benchmark result.

Inkling is a 975-billion-parameter mixture-of-experts model trained on forty-five trillion tokens across text, image, audio, and video, with a context window stretching to a million tokens. That is frontier-scale compute expenditure. This isn’t a company that ran out of runway and settled for a smaller ambition. It’s a company that spent frontier-level resources and then declined to spend the final increment chasing benchmark supremacy, presumably because the return on that increment doesn’t show up in the business they’re building.

A second detail: Inkling reportedly uses one-third the tokens of Nemotron 3 Ultra to hit equivalent performance on agentic coding benchmarks. That’s not a capability retreat — that’s a capability choice, optimizing for efficiency and cost-per-task rather than raw benchmark position. And the company is previewing a smaller sibling model alongside Inkling, suggesting a family strategy across sizes rather than a single mid-tier release.

The restaurant next door: Palantir

Thinking Machines isn’t the only company making this bet — and looking at who else is making it shows not everyone is occupying the same layer.

Earlier this month Palantir and Nvidia announced a “Sovereign AI Operating System” — Nvidia’s open Nemotron models, fine-tuned on a customer’s own data, running on Nvidia hardware inside that customer’s own air-gapped network, with Palantir’s Ontology and Foundry software layered on top. CEO Alex Karp pointed out that his enterprise customers don’t want to risk sharing their IP with frontier model providers and asked simply why wouldn’t they control the weights?

It’s tempting to read this as the same argument Thinking Machines is making. It isn’t, quite. What Palantir is selling is the restaurant: the finished, seasoned, plated product — an air-gapped AI system wired into a specific government agency’s or enterprise customer’s actual workflows, with “you control the weights” as the pitch that closes the deal. Palantir isn’t growing wheat. It’s the chef, working with ingredients somebody else grew. Somebody who could be trusted.

Another restaurant: Sierra

Sierra, Bret Taylor and Clay Bavor’s customer-support agent company, makes the same choice even more starkly. Sierra’s own technical writing describes a “constellation of models” architecture: rather than betting on a single LLM, Sierra routes each task inside a customer-service agent to whichever model — from OpenAI, Anthropic, Meta, or elsewhere — handles it best, and explicitly says it invests “in fine-tuned models where off-the-shelf models fail to meet our constraints.” Fine-tuning shows up in Sierra’s stack as one tool among several, alongside retrieval and layered “supervisor” models that catch mistakes before a customer sees them. Sierra has no interest in being a model company or an infrastructure company. It wants to be the restaurant that happens to keep a few specialty ingredients in the walk-in that nobody else stocks, because the dish needs them and they know just how to include them.

Mapping the rest of the stack

The farm-kitchen-restaurant split shows up everywhere the fine-tuning economy has organized itself.

The kitchen-builders — companies selling fine-tuning infrastructure to whoever wants to cook with it, indifferent to what gets made — now form a crowded field: Thinking Machines’ Tinker, Together AI, Fireworks AI, Predibase, OpenPipe, Baseten, Modal, Databricks’ Mosaic stack, and newer entrants like Nebius’s Token Factory and Prime Intellect. None of them care whether you’re building a coding agent, a legal research tool, or a customer-service bot.

The restaurants — companies where fine-tuning is invisible plumbing inside a finished, vertical product — include Palantir and Sierra, and many others. The addressable market for fine-tuning seems to include almost every possible enterprise adopting AI.

What’s seems unusual about Thinking Machines is that it’s trying to be the farm and the kitchen simultaneously while declining, by its own public admission, to be the restaurant. Most companies pick one layer and defend it. Thinking Machines is betting that owning two of the three is the more durable position — grow the flour, own the stove, and let Palantir, Sierra, and a thousand enterprise engineering teams fight over who plates the dish.

The same stack, built by design

As I was thinking about this, I wondered how this relates to the AI activities underway in China. It seems that China’s AI industry maps onto this same three-layer structure with unusual clarity — and one genuine wrinkle the American version doesn’t have.

The farm is crowded and innovating on a different axis than size: DeepSeek, Alibaba’s Qwen, Zhipu AI, Moonshot AI, MiniMax, ByteDance’s Doubao and Seedance. The standout isn’t scale, it’s efficiency — DeepSeek’s V3.2 reportedly uses a novel sparse attention mechanism to nearly match GPT-5 and Gemini 3 on complex reasoning despite far less compute, a different kind of farming: not more wheat, but wheat bred to need less water. Qwen has become the default soil for the rest of the world’s kitchens, generating over 100,000 derivative fine-tunes on Hugging Face. VC’s in Silicon Valley note how frequently their startup companies are building on Qwen.

The kitchen layer has its own SiliconFlow — a Beijing infrastructure startup, backed by Alibaba Cloud, that bills itself as the neutral layer between AI applications and hardware. It solves a problem others never had to: China’s compute runs across fragmented domestic chips, Huawei’s Ascend line chief among them, that don’t share Nvidia’s CUDA ecosystem. SiliconFlow abstracts that fragmentation away — it became the fastest platform serving DeepSeek traffic, and the only large provider running DeepSeek on Ascend chips instead of Nvidia’s. That’s a stove engineered to burn whatever fuel is in the tank that week, a direct product of the U.S. chip export controls rather than any inherent technical edge. Volcano Engine, Alibaba Cloud’s PAI, and Baidu’s Qianfan are versions of the same layer.

The restaurant layer is where China’s picture diverges most from Palantir and Sierra’s venture-funded improvisation: it’s named industrial policy.

Beijing’s “AI+” initiative targets seventy percent sectoral AI penetration by 2027, ninety by 2030 — fine-tuned vertical deployment treated the way past five-year plans treated high-speed rail. The players read like a sector directory: SenseTime for vision and embodied AI, iFlytek for speech in education and government, Baichuan Intelligence for healthcare, 4Paradigm for finance and industry, each fine-tuning a general base into something that only makes sense inside one workflow — a hospital’s diagnostic support tool, a bank’s risk model, an industrial inspection line.

The bet

Every frontier lab is implicitly betting that intelligence is the scarce resource, and that whoever has the most of it wins the enterprise market by default. Thinking Machines, Palantir, Sierra, and many others are all, in their different ways, betting against that premise — that raw intelligence is commoditizing faster than the frontier labs’ spending would suggest, and that the scarce resource has already migrated to whichever layer turns a generalist model into a specific customer’s model.

Thinking Machines is betting the moat moved to the farm-and-kitchen layer. Palantir, Sierra and others are betting it moved further still, to the restaurant, where nobody cares whose flour was used as long as the dish is right. China is betting on all three layers at once, with the state underwriting the bet directly.

It is a curious thing for me to watch companies with this much money and this much talent choose not to fight for the title of smartest model in the room. It is also a curious thing to watch them explain why, in public, in the first paragraph of an announcement.

But I think I’m beginning to understand.

Categories
AI China

Model Competition

The Chinese AI company Moonshot AI released its Kimi K3 model yesterday. I played with it a bit in the Kimi iOS app and, frankly, was underwhelmed. It was slow and disappointing. But this was hardly an extensive test. It’s strictly my first impression. I’m not doing coding. I’m what model makers would call a knowledge worker.

This morning Gavin Baker posted on X about Kimi K3 calling it potentially an “important inflection point for AI”:

Kimi K3 may be an important inflection point for AI. Potentially negative for Anthropic and OpenAI while being net positive for essentially every other company in the world. I mean that very literally. Although the real “Sputnik moment” would be an open-source frontier model that was also token efficient unlike Kimi K3 which is 50-70% more expensive to run than GPT 5.6 per Artificial Analysis.

Rationale:
 
A world where there are only 2-3 dominant frontier labs with 90% inference margins is net negative for every other layer while being awesome for those 2-3 labs. Those labs would become monopsonies for power, data centers, semiconductors and hyperscalers and would obviously vertically integrate over time into all those layers while also completely subsuming the application/software layers. 
 
Anything that lowers margins and increases competition at the model layer is good for every other AI layer: power, semiconductors, hyperscalers, neoclouds and yes even software.
 
This is why Jensen is so supportive of open-source. An open-source model requires the exact same amount of compute to run as a closed frontier model of similar size and architecture. Kimi K3 is roughly the same price as GPT 5.6 Terra on a per token basis, which actually suggests that it is less computationally efficient as I am sure that GPT 5.6 is priced to a higher margin than K3. And given that K3 is a token wastrel, i.e. token inefficient, it is significantly more expensive per task than GPT 5.6 and Grok 4.5, which are much more token efficient. Cost per token and token efficiency (i.e. intelligence density per token) are the drivers of intelligence per unit of cost. The winning AI companies will be those that offer the most intelligence per $ over time.
 
Lower margin % at the model layer = more margin $ at every part of the infrastructure layer and is a godsend for software. This can happen either through open-source models like K3 at the frontier or having a vertically integrated model company like Meta, SpaceX or Google at the frontier. Both outcomes result in a lower margin % at the model layer as vertically integrated model companies don’t really care where the margin $ come from. This is why it was so painful for OpenAI and Anthropic when Google was right there with them from a model competitiveness perspective and why Grok 4.5 and Muse 1.1 were just as important as Kimi K3.

The reason Kimi K3 is only potentially negative for Anthropic and OpenAI is 1) the @ericvishria point that the Claude and ChatGPT products and harnesses may be more important than their models today and 2) the hypothesis that they have much more advanced model checkpoints internally that are already being used for RSI. In the latter scenario, reaching RSI even a few months ahead of other labs might be enough to cement a permanent lead.

Time will tell on both points. And likely fairly quickly.

Caveat would be that since Kimi K3 is not token efficient and thereby actually more expensive than ChatGPT 5.6, we may need to see a more token efficient open-source model at the frontier or see Grok 5/Composer 4/Muse 2 at multiple points on the Pareto frontier for this potential risk to Anthropic and OpenAI to play out. And I am sure they will both vertically integrate as quickly as possible while continuing the product/harness strength they have shown over the last 8 months.

I wanted to explore this further so I asked Grok to explain his post:

This is a tweet (X post) from Gavin Baker (@GavinSBaker), Managing Partner and CIO of Atreides Management, posted on July 17, 2026.

It analyzes the recent launch of Kimi K3 (by Chinese AI company Moonshot AI) as a potentially significant shift in the AI industry. Baker argues it could hurt the high-margin dominance of closed U.S. frontier labs like OpenAI and Anthropic while benefiting nearly everyone else in the broader AI ecosystem (infrastructure, chips, data centers, power, hyperscalers, and software).

Quick Context on Kimi K3

Kimi K3 is Moonshot AI’s new flagship model (launched ~July 16, 2026). Key specs include:

  • ~2.8 trillion parameter Mixture-of-Experts (MoE) architecture.
  • Massive 1 million token context window.
  • Native vision capabilities.
  • Strong early performance in coding, agentic tasks, and long-context reasoning (competitive with or close to top closed models like GPT-5.6 variants or Claude models in certain benchmarks).
  • Pricing: $3 input / $15 output per million tokens (with caching discounts).
  • Important: Full weights are scheduled for open release (modified MIT license) around July 27, 2026 — making it one of the first truly open frontier-scale models.

It represents a step up from prior Kimi models and adds meaningful competition from a non-U.S. lab.1

The Two Images in the Tweet

These are screenshots from Artificial Analysis (an independent benchmarking site):

  1. Bar chart — “Cost per Intelligence Index Task” (lower is better). It breaks down costs by token type (input, output, reasoning, cache, etc.) and shows Kimi K3 as relatively expensive compared to some GPT-5.6 variants.
  2. Scatter plot — Intelligence Index (y-axis) vs. Cost per task (x-axis, log scale). It positions models in quadrants, with the “most attractive” area being high intelligence at low cost. Kimi K3 appears competitive on intelligence but not the cheapest.

These visuals support Baker’s point that while Kimi K3 reaches frontier-level intelligence, it is not the most token-efficient, making its effective cost per useful output higher than more efficient models.

Baker’s Core Argument (Broken Down)

Baker’s thesis is economic and structural, not just technical:

  • A world with only 2–3 dominant closed frontier labs (with ~90% inference margins) is bad for almost everyone else.
    Those labs would act as monopsonies (dominant buyers) for power, chips, data centers, and compute. They would vertically integrate upward (into infrastructure) and downward (into apps/software), capturing most of the value. Competition at the model layer prevents this concentration.
  • Anything that increases competition or lowers margins at the model layer is broadly positive.
    It redistributes economic value to the rest of the stack: more demand for GPUs/chips (inference especially), power, data centers, hyperscalers/neoclouds, and software (because cheaper/better AI drives wider adoption).
  • Token efficiency matters more than raw price per token.
    Baker notes Kimi K3 is roughly the same per-token price as GPT-5.6 but 50-70% more expensive overall according to Artificial Analysis. It is a “token wastrel” — it uses more tokens to achieve similar results (lower “intelligence density per token”).
    → True winners deliver the most intelligence per dollar. Cost-per-token + efficiency together determine real value.
  • Why NVIDIA’s Jensen Huang loves open-source models.
    Open models still require massive compute to run at scale (same as closed ones). More models + wider adoption = more inference demand = more GPUs sold. A few ultra-high-margin closed labs would concentrate spend and reduce overall volume.
  • Vertical integration changes the game.
    Companies like Google, Meta, or xAI/SpaceX that have both strong models and their own infrastructure don’t obsess over model-layer margins — they capture value elsewhere. This is why more competition (open or vertically integrated) hurts pure-play high-margin model companies like OpenAI/Anthropic.
Why It’s Only “Potentially” Negative for OpenAI and Anthropic

Baker is careful:

  • Their products, user interfaces, and “harnesses” (the surrounding tools/ecosystems) may matter more than raw model performance right now.
  • They likely have more advanced internal models/checkpoints (possibly already used for recursive self-improvement/RSI).
  • Kimi K3’s token inefficiency softens the blow for now. A more efficient open-source frontier model (or stronger showings from Grok, etc.) would be more threatening.
  • OpenAI and Anthropic are already moving to vertically integrate.
Overall Takeaway from the Post

Baker sees Kimi K3 (and the wave of strong Chinese models like Qwen, DeepSeek, GLM) as healthy competition that prevents unhealthy concentration. It’s not yet a full “Sputnik moment” (a truly dominant, efficient open-source frontier model), but it points in that direction.

Net effect he describes: Lower % margins at the model layer → more total margin dollars flowing to infrastructure, chips, power, and software layers. This is bullish for the broader AI ecosystem even if it pressures the pure closed frontier labs.

The tone is analytical and balanced — he acknowledges uncertainties and that time will tell (likely quickly). Replies to the post generally agree this kind of competition is positive for the industry overall and could shift more spend toward efficient inference hardware.

This fits Baker’s broader investing views on AI: he often emphasizes infrastructure winners, token economics, and how value accrues across the full stack rather than just at the smartest model.

Interesting.

Categories
AI

Memory Optimization

After thinking about the very high margins that Micron recently reported (greater than Nvidia’s) I asked Grok 4.5:

What startups are pursuing technology that reduces memory requirements for ai models and inference?

Several startups are actively developing technologies to slash memory requirements for AI models and inference. These efforts target the core bottlenecks in large language models (LLMs) and multimodal systems—primarily through extreme quantization (reducing bit precision of weights), alternative architectures that avoid quadratic memory scaling (e.g., state-space models or hybrids), sparsity/partial activation (like Mixture-of-Experts or neuron-level sparsity), efficient runtimes/kernels, and hardware innovations like in-memory computing that minimize data movement.

This space is especially active for on-device/edge AI (phones, laptops, embedded systems) and cost/energy-efficient cloud inference, where fitting larger or more capable models into limited RAM/VRAM, reducing KV cache bloat in Transformers, or lowering bandwidth demands delivers major gains.

Here are some of the most prominent startups and their approaches (as of mid-2026):

Extreme Quantization & Model Compression

  • PrismML (Caltech spinout, backed by Khosla Ventures): Focuses on extreme low-bit architectures—1-bit binary (“Bonsai”) and ternary (1.58-bit, weights in {-1, 0, +1}). They compressed Alibaba’s Qwen3.6 27B model from ~54 GB (FP16) to under 4 GB and ran the full dense 27B model on an iPhone 17 Pro. Claims include up to 14× smaller memory footprint, 8× faster inference, and significantly lower energy use, with competitive or better benchmark performance. They have open-sourced Bonsai models (including smaller 8B/4B/1.7B variants) under Apache 2.0 and are in discussions with Apple. This represents one of the most aggressive commercial pushes into 1-bit/ternary models for on-device deployment.
  • Mobius Labs (Berlin): Developed Half-Quadratic Quantization (HQQ), a fast, calibration-light post-training quantization method that enables high-accuracy low-bit models (including aggressive 2-4 bit). They demonstrated quantizing Llama 70B to run on a single GPU instead of four without major accuracy loss, directly cutting memory and compute needs. Their work extends to FP4 optimizations and integrates with frameworks like vLLM.
Alternative Architectures for Inherent Memory Efficiency
  • Liquid AI (MIT spinoff): Builds Liquid Foundation Models (LFM / LFM2 series)—hybrid architectures combining gated short convolutions with grouped-query attention (GQA) blocks, plus MoE variants. These deliver substantially lower memory footprints than Transformers (especially for long contexts, avoiding massive KV cache growth), faster prefill/decode (up to 2× on CPU in some cases), and strong on-device performance. Examples include tiny models (230M–350M params, often
  • Cartesia: Specializes in state-space models (SSMs) and Mamba-derived/hybrid architectures. These offer near-linear (or better) scaling with sequence length and near-constant memory complexity during inference (fixed or minimal growth vs. quadratic KV cache in Transformers). This enables efficient long-context processing on smaller devices with lower energy use. They provide on-device libraries (e.g., “Edge”) and models like hybrid Mamba-2 variants optimized for throughput and fixed memory footprints.
On-Device/Edge Runtimes & Optimizations
  • Cactus (Y Combinator-backed): Builds cross-platform, energy-efficient inference runtimes and kernels for mobile, wearables, and low-power devices. Supports aggressive quantization (down to 2-bit), delivers very low latency (sub-50 ms time-to-first-token), and keeps models small (e.g., hundreds of MB for small models). Emphasizes privacy (fully local) and broad model support (Llama, Qwen, Gemma, etc.).
  • Embedl: Focuses on edge-optimized multimodal models and mixed-precision quantization (e.g., W4A16-Edge2 with selective FP16 layers + “FlashHead” optimizations). Enables fast multimodal reasoning on devices with under 8 GB RAM (e.g., Jetson Orin Nano/AGX), with significant speedups over baselines while recovering most accuracy.
Hardware Innovations Reducing Memory Pressure
  • d-Matrix: Develops digital in-memory computing (DIMC / 3DIMC) chips and platforms (e.g., Corsair) that integrate compute directly with/near memory (initially SRAM, now advancing to 3D DRAM). This drastically cuts data movement—the dominant memory bandwidth and energy bottleneck in inference. Claims include major speedups (up to 10× in some workloads) and better efficiency vs. traditional GPU/HBM setups, enabling higher performance with more manageable overall memory system requirements.
  • Hailo: Designs efficient edge AI inference chips (e.g., Hailo-10H series) optimized for low-power LLM and multimodal workloads. Their chips target 1–3B parameter models at very low power envelopes (~2–2.5 W), making deployment feasible on memory- and power-constrained edge devices where traditional hardware would struggle.
Other Notable Mentions
  • Tiiny AI: Uses neuron-level sparse activation (“TurboSparse”) + heterogeneous engines for running very large models (up to 120B) fully on-device/offline in compact hardware.
  • Broader ecosystem players (inference platforms like Together AI or Fireworks) heavily optimize serving (quantization, caching, KV cache management) but are more infrastructure-focused than core technology innovators in compression/architecture.
Key Trends and Context

Quantization (especially post-training or quantization-aware) remains the most immediate lever and is widely adopted. Architectural shifts (SSMs, hybrids, MoE with low active parameters, convolutions) address fundamental scaling issues like KV cache memory growth. Hardware approaches (in-memory compute, high on-chip SRAM, efficient NPUs) tackle bandwidth and data movement. Many solutions combine these (e.g., architecture + quantization + sparsity).

Progress is rapid, driven by on-device AI demands (Apple’s interest in PrismML is notable) and the economics of scaling inference. Many models are open-sourced or have open weights, accelerating adoption. Results vary by workload—extreme 1-bit models excel in certain scenarios but may need custom kernels; SSM/hybrid models shine for long context or linear scaling needs.

For the absolute latest benchmarks or specific model deployments, check company sites (e.g., PrismML’s Bonsai releases, Liquid AI’s LFM models on Hugging Face) or recent evaluations, as the field evolves quickly.

Interesting!

Categories
AI Business

The Reverse Information Paradox We’ve Always Had

Satya Nadella wrote recently about what he calls the Reverse Information Paradox: enterprises pay for AI intelligence twice. Once in money. Again in the proprietary knowledge they surrender through every prompt, correction, and evaluation. The better they use the model, the more of their own institutional understanding leaks into someone else’s system. The vendor ends up knowing more about the buyer’s business than the buyer knows about what the vendor retained.

Replace “model” with “employee” (or “consultant”) and the paradox is not new at all.

You pay for a person once with salary. You pay again with something harder to price: the context, relationships, and judgment they must absorb to become useful to you. The better they perform, the deeper the immersion, the more of your particular way of doing things moves into their head. Every correction and late-night conversation is another trace of institutional memory changing hands. When they leave, some of that memory leaves with them. Not always through theft. Usually just through the ordinary residue of good work.

The visible cost is salary; the invisible cost is the slow transfer of what makes you distinctive. High performers get more access precisely because they’re high performers, which means the leakage accelerates exactly when you can least afford it. The exhaust is just harder to see with people than with tokens — it moves through conversation and mental models instead of logs.

The analogy has a limit, and the limit matters. Employees bring knowledge in, not just absorb it. They have judgment and relationships a model doesn’t. Models are purely absorptive, and once something is inside them, it’s infinitely reproducible — a person can only be in one place, working for one employer, at a time. We’ve had a few hundred years to build tools for the human version of this problem: contracts, culture, non-competes. The model equivalent is still being invented in real time, which is exactly why Nadella felt the need to name it.

Apple’s recent legal action against former employees who joined OpenAI is this pattern in its sharpest form. Whatever the specifics, the shape is familiar: people who spent years inside one of the most sophisticated organizations in the world, carrying out knowledge that never appeared on any balance sheet and was hard to contain. No one fully anticipates what a mind absorbs simply by being in the room long enough.

That’s the real difference between the silicon case and the human one. You can try to take action to wall off knowledge flowing to a model. You cannot wall off what someone has learned to notice.

Categories
AI Apple Google

The Library You Already Own

Sharon Park in the morning is not a dramatic place. There’s a duck pond, a stand of oaks that go gold too briefly in November, and a loop I’ve walked enough times that my legs know it better than my eyes do. It is, in other words, exactly the kind of place where a person starts talking to himself. Not out loud. In the productive, low-grade way — turning a sentence over, arguing with an idea from the day before, checking a thought against something you believe about yourself.

I think in five years I’ll be doing that walk with something else along. Not a search engine. Not another chatbot trained to know a little about everything and a lot about nothing in particular. Something closer to a second set of eyes on my own life — a reasoning engine, lean and mostly private, that has actually read the things I’ve written and doesn’t need me to explain who I am before it’s useful.

Here’s the distinction that matters, and it took me longer than it should have to see it clearly. The AI industry has spent years in an arms race over how much of the world a model can hold — more facts, more languages, more of the internet compressed into weights. That race will keep going, and somebody else can have it. What I want is smaller and stranger: a model that knows comparatively little about the world and quite a lot about me. My core values document. The portfolio spreadsheets. Fifteen years of blog posts. The half-finished notes for the I-280 project, sitting in a folder, waiting for someone — or something — to ask the right question about them.

I spent a career in payments infrastructure, which means I spent a career thinking about a very specific kind of trust: the kind where a stranger’s system has to make a judgment call, in milliseconds, about whether to say yes. Fraud models don’t work because they know everything about commerce. They work because they know an enormous amount about one account, one pattern, one person’s ordinary Tuesday — enough to notice when Tuesday stops being ordinary. That’s the architecture I keep picturing, aimed inward instead of outward. Not a system trying to know the world. A system trying to know me, well enough to notice when I’m drifting from what I said I cared about.

I can already feel the shape of the mornings this would change. Right now, when I sit down to look at RMD requirements against the tax picture, I’m doing the translation myself — pulling numbers into a story I can actually feel the weight of. A reasoning engine grounded in my real holdings wouldn’t just run the scenario. It would know that I don’t want the scenario dressed up as a spreadsheet; I want it dressed up as a conversation, unhurried, the kind you’d have over lunch with someone who already knows the whole situation. And on the mornings when I sit down to write, instead of staring at a blinking cursor and a blank page that has no idea I exist, I’d be handing a draft to something that has actually read my last two hundred posts and knows the difference between the sentence I’d write and the sentence I’d cut.

None of this is especially exotic technology. Apple and Google are already building toward it — Neural Engines fast enough to do real reasoning on-device, retrieval systems that can reach into your own files instead of the entire internet, fine-tuning that’s getting cheap enough to personalize rather than merely customize. The more interesting story here isn’t privacy, though privacy is real. It’s architectural: what happens when the expensive, impressive part of the system — the part that knows everything — becomes optional, and the cheap, personal part — the part that knows you — becomes the whole point.

What I don’t yet know is what this will cost me. A tool that reasons this well about my own life is also a tool I could lean on instead of doing the leaning myself, and there’s a version of this future where the walk around Sharon Park stops being mine and starts being a conversation with something that finishes my sentences a little too well. I’d want some way of knowing, plainly, what it’s drawing from and what it’s guessing at — less a nutrition label than a kind of honesty I could check against, the way you’d check a fraud model’s confidence score before you trusted it with a yes.

But most mornings, I think I’d take the trade. Not because I want to think less. Because for thirty years I’ve been collecting the raw material — the notebooks, the portfolios, the half-built essays — and it would be something, finally, to walk beside a mind that had actually done the reading.

Categories
AI

The Encyclopedia and the Reasoner

I was standing in the cereal aisle a few weeks ago, doing the thing I always do — flipping the box over, scanning the fine print, comparing fiber grams like it mattered more than it probably does — when I thought about the model I’d been testing that morning. Sharp. Fast. Occasionally, confidently, wrong about something I could have looked up in ten seconds.

There was no label for that. No panel telling me what was inside, what it was good at, what it might get wrong, what it cost to run. Just a chat window and a kind of blind trust.

That’s the itch behind this post. What would it look like if AI models came with something like a Nutrition Facts label — the kind the FDA forced onto every box in your pantry back in 1994? Not as a gimmick, but as a real answer to a real problem: we are feeding these things into our decisions, our writing, our portfolios, our kids’ homework, largely on faith.

The IQ Number That Isn’t Quite an IQ Number

I keep running into a shorthand in investing circles — Jordi Visser and others talking about frontier models as “140 IQ” systems, reasoning at a level that outpaces most humans on the kinds of puzzles we associate with fluid intelligence. Pattern recognition. Logic chains. Novel deduction under pressure.

It’s a useful number. It’s also a bit of a trick.

Human IQ tests were built to measure something narrow and specific — not wisdom, not knowledge, not judgment, but the raw machinery of reasoning. When we borrow that language for AI, we inherit the same narrowness, which is fine as long as we remember it. A model that aces abstract reasoning benchmarks isn’t necessarily the model that knows the correct dosage, the right case law, or what actually happened in 1932. Reasoning and knowledge are cousins, not twins.

Two Kinds of Smart

Here’s an old-fashioned way to think about the split: Britannica versus World Book.

Britannica was the encyclopedia my father would have trusted — dense, expert-written, unapologetically deep, assuming you could keep up. World Book was the one actually sitting on the shelf in most houses I knew growing up, mine included: friendlier, broader, built for a general reader, a little shallower in exchange for being a little more useful on a Tuesday night with a homework assignment due.

Neither is wrong. They’re optimized for different things. And training data does the same kind of sorting. A model fed heavily on curated, scholarly, expert-vetted sources leans Britannica — deep, careful, occasionally slow to update. A model trained on the sprawl of the open web leans World Book — broad, current, occasionally sloppy, sometimes brilliant at the edges precisely because it’s seen everything.

Any honest label for a model needs a section on this. Call it “Knowledge Sourcing.” Not just how big the training set was, but what kind of encyclopedia it’s pretending to be.

Sketching the Label

If I could design the box myself, it might read something like this:

Serving Size: 1 query, ~500 tokens

Reasoning Score: 138 (fluid problem-solving, logic, abstraction) Knowledge Depth: Moderate–High (cutoff: [date]; strongest in [domains]; weakest in [domains])
Ingredients: Curated scholarly corpora, licensed news archives, public web crawl, synthetic reasoning data, human feedback Allergens: Confident hallucination under ambiguous prompts; recency gaps beyond training cutoff; known weakness in [specific domain]
Cost per Serving: $X per million tokens; Y watt-hours per query Best Paired With: Retrieval tools, human review for high-stakes decisions

It’s a little tongue-in-cheek written out like that. But underneath the joke is something I actually want — the same instinct that made me read cereal boxes as a kid. Not to be scared of what’s inside, just to know.

The Part That Actually Excites Me

Here’s where the scaling laws get interesting, and where I think the real opportunity sits.

World knowledge is expensive. It’s greedy for data and parameters — you need to have practically read the internet to know the boiling point of tungsten, the plot of a minor Victorian novel, and the org chart of a mid-cap company all at once. Reasoning, it turns out, is a different kind of animal. It can be distilled, compressed, taught through synthetic problems and careful post-training, and squeezed into something far smaller than you’d expect.

Which means a genuinely thrilling possibility is already taking shape: sharp, high-reasoning models small enough to run on a phone or a laptop, entirely offline, because they’ve shed the encyclopedia and kept the mind. Pair one of those with a personal index — your own notes, your own documents, a retrieval layer built around your actual life — and you get something closer to a personal thinking partner than a general-purpose oracle. Private. Fast. Always available. Tuned to you rather than to everyone. Apple may be on to something with this kind of strategy?

I think about this constantly in my own workflow — the daily scans, the little agents I’ve built to help sort signal from noise, the genealogy digging, the investment frameworks I keep refining. What I usually want isn’t more encyclopedia. It’s a clear-headed reasoner sitting next to my own carefully kept knowledge, not buried under someone else’s version of the whole internet.

Why the Label Matters More Than the Score

None of this works, though, without honesty about what’s inside the box. A 140 on a reasoning benchmark tells you almost nothing about whether a model will quietly misremember a fact it was never that confident about in the first place. And a model can be extraordinarily knowledgeable while being a mediocre reasoner — plenty capable of reciting the right ingredients and still getting the recipe wrong.

The nutrition label movement in food didn’t eliminate junk food. It just made it possible to choose junk food on purpose, with your eyes open, instead of by accident. I’d like the same deal with AI. Not a demand that every model be a genius generalist, but a demand that I get to know what I’m actually consuming — and choose the lean local thinker over the bloated encyclopedia when that’s what the moment calls for, or the other way around when it isn’t.

Curiosity got me into that cereal aisle habit decades ago, and it’s the same instinct pulling me toward this idea now — not suspicion of the box, just a wish to read it clearly before I decide how much of it to trust.

What would you want on your label?

Categories
AI Semiconductors

The Margin of the Weather

A company that has sold memory chips for forty years — memory, one of the most humiliatingly commoditized products in capitalism, a business that has bankrupted entire Korean and Japanese conglomerates teaching each other lessons about discipline — is about to make more money in twelve months than in the previous four decades combined.

Samsung’s chip chief told a room of his own employees: this year’s profit will exceed everything the division has earned since the 1970s. Forty years of grinding, erased by one fiscal year. You’d think they’d invented something.

They hadn’t. Everyone building an AI data center needs memory. Nobody built enough factories. Samsung was one of three companies on earth able to supply the shortfall, and the price of a chip that costs what it always cost went up fifty percent. Samsung kept the difference. Not innovation. What happens to a farmer when the drought hits every field but his.

We don’t credit the lucky farmer with genius. We say: good year. And we don’t expect the good year to repeat. Rain comes back. The price falls. Scarcity is weather, not a personality trait.

There’s a real achievement in this story too, and it has nothing to do with the weather. A year ago Samsung failed to qualify its most advanced memory for Nvidia’s systems — performance problems, a rival getting the business instead. The engineers went back and fixed it. That’s the actual skill in this company’s year: unglamorous, uncelebrated at the town hall, worth nothing next to the number that got the confetti. The competence arrived quietly, on a different chip, in a different meeting, and nobody’s putting that on a plaque.

The stock market didn’t put it on one either, but it seemed to know the difference. Best quarter in Samsung’s history — profit nineteen times the year before — and the shares fell seven percent. Not despite the earnings. The gain had already been priced in, the shares having run up a hundred and fifty percent on the expectation of exactly this number, so the number’s arrival became a ceiling instead of a floor. A market rewards discovery. It does not reward weather. Had investors believed Samsung built something durable — the Nvidia qualification, the years of engineering behind it — the stock would have ripped, the way See’s Candies or Apple gets rewarded quarter after quarter, because everyone agrees the thing generating the money isn’t going anywhere. Instead the market glanced at the record harvest and asked, politely, whether it would rain again next year.

Analysts insist the shortage holds through next year. Someone always insists that, right before it doesn’t. Fabs get built. Capacity catches the demand that summoned it, the way it always has, and the cycle ends the way memory cycles end — too much supply chasing too little demand, margins reverting toward the number they were always going to revert toward. Nobody knows if this time is different. A company just posted the best year of its life, on a windfall it didn’t earn and a fix it did, and the market — which has seen droughts end before — hasn’t decided yet which one it’s watching.