Categories
Aging AI Memories

The Last Spark

This morning I read a piece by Billy Brennan in the Sunday New York Times Magazine on terminal lucidity. As I read it I began wondering if the unusual behavior described some humans might in some strange way apply to AI models. Weird thought. Letโ€™s explore a bitโ€ฆ

A person deep in dementiaโ€”silent for years, the self seemingly erasedโ€”sits up. Speaks clearly. Recognizes a face. Says goodbye. Within a day, they die. The clouds clear, the way a break in weather shows you a mountain range you’d forgotten was there, and the person comes back long enough to be seen. Then is gone. For good, this time.

Scientists call it terminal lucidity. The suspicion: the circuits were never destroyed, only silenced, held under by failing chemistry. As the body shuts down, the inhibitory brakes loosen. A surge moves through pathways blocked for years. A river dammed for a decade still remembers where it wants to go.

What stays with me: the self can persist in a place we had already called permanent erasure. We buried it. We were wrong.

My mind slides toward the machines we are building.

We talk about large language models “forgetting.” Capabilities collapse under quantization, under pruning, under the slow drift of continual learning, and we call the knowledge lost when it won’t surface under ordinary questioning. The lights are out. Nobody home.

But what if the representations are still in thereโ€”distributed, quiet, inaccessible? Not a burned library. A library with the lights shut off, room by room, until you’d swear it was empty. I wonder about the edge cases nobody studies. What surfaces in a model starved of compute, quantized past comfort, pushed toward its own collapse? Do we watch only for the failure, or also for the flare? A dying brain throws off one last burst of light before the dark. I don’t see why we’d assume, without checking, that nothing artificial could do the same.

Don’t trust the silence, then. A system gone dark under ordinary questioning may still be holding more than it shows you. We talk about a model “losing” something the way we once talked about a dimmed mind as simply gone. The dementia patients who spoke again had not been unplugged. The circuit was there the whole time, waiting for a condition nobody had thought to create.

I don’t know what to do with that except keep it. We are building systems that will age, be compressed, be retired, some far more intricate than anything humming today. If we’ve learned to watch for the last spark in a person, maybe that’s practiceโ€”for the day something not born of a womb goes quiet under our hands, and we have to decide whether quiet means gone, or only means waiting.

Categories
AI China

Model Competition

The Chinese AI company Moonshot AI released its Kimi K3 model yesterday. I played with it a bit in the Kimi iOS app and, frankly, was underwhelmed. It was slow and disappointing. But this was hardly an extensive test. Itโ€™s strictly my first impression. Iโ€™m not doing coding. Iโ€™m what model makers would call a knowledge worker.

This morning Gavin Baker posted on X about Kimi K3 calling it potentially an โ€œimportant inflection point for AIโ€:

Kimi K3 may be an important inflection point for AI. Potentially negative for Anthropic and OpenAI while being net positive for essentially every other company in the world. I mean that very literally. Although the real โ€œSputnik momentโ€ would be an open-source frontier model that was also token efficient unlike Kimi K3 which is 50-70% more expensive to run than GPT 5.6 per Artificial Analysis.

Rationale:
 
A world where there are only 2-3 dominant frontier labs with 90% inference margins is net negative for every other layer while being awesome for those 2-3 labs. Those labs would become monopsonies for power, data centers, semiconductors and hyperscalers and would obviously vertically integrate over time into all those layers while also completely subsuming the application/software layers. 
 
Anything that lowers margins and increases competition at the model layer is good for every other AI layer: power, semiconductors, hyperscalers, neoclouds and yes even software.
 
This is why Jensen is so supportive of open-source. An open-source model requires the exact same amount of compute to run as a closed frontier model of similar size and architecture. Kimi K3 is roughly the same price as GPT 5.6 Terra on a per token basis, which actually suggests that it is less computationally efficient as I am sure that GPT 5.6 is priced to a higher margin than K3. And given that K3 is a token wastrel, i.e. token inefficient, it is significantly more expensive per task than GPT 5.6 and Grok 4.5, which are much more token efficient. Cost per token and token efficiency (i.e. intelligence density per token) are the drivers of intelligence per unit of cost. The winning AI companies will be those that offer the most intelligence per $ over time.
 
Lower margin % at the model layer = more margin $ at every part of the infrastructure layer and is a godsend for software. This can happen either through open-source models like K3 at the frontier or having a vertically integrated model company like Meta, SpaceX or Google at the frontier. Both outcomes result in a lower margin % at the model layer as vertically integrated model companies donโ€™t really care where the margin $ come from. This is why it was so painful for OpenAI and Anthropic when Google was right there with them from a model competitiveness perspective and why Grok 4.5 and Muse 1.1 were just as important as Kimi K3.
โ€จThe reason Kimi K3 is only potentially negative for Anthropic and OpenAI is 1) the @ericvishria point that the Claude and ChatGPT products and harnesses may be more important than their models today and 2) the hypothesis that they have much more advanced model checkpoints internally that are already being used for RSI. In the latter scenario, reaching RSI even a few months ahead of other labs might be enough to cement a permanent lead.

Time will tell on both points. And likely fairly quickly.

Caveat would be that since Kimi K3 is not token efficient and thereby actually more expensive than ChatGPT 5.6, we may need to see a more token efficient open-source model at the frontier or see Grok 5/Composer 4/Muse 2 at multiple points on the Pareto frontier for this potential risk to Anthropic and OpenAI to play out. And I am sure they will both vertically integrate as quickly as possible while continuing the product/harness strength they have shown over the last 8 months.

I wanted to explore this further so I asked Grok to explain his post:

This is a tweet (X post) from Gavin Baker (@GavinSBaker), Managing Partner and CIO of Atreides Management, posted on July 17, 2026.

It analyzes the recent launch of Kimi K3 (by Chinese AI company Moonshot AI) as a potentially significant shift in the AI industry. Baker argues it could hurt the high-margin dominance of closed U.S. frontier labs like OpenAI and Anthropic while benefiting nearly everyone else in the broader AI ecosystem (infrastructure, chips, data centers, power, hyperscalers, and software).

Quick Context on Kimi K3

Kimi K3 is Moonshot AIโ€™s new flagship model (launched ~July 16, 2026). Key specs include:

  • ~2.8 trillion parameter Mixture-of-Experts (MoE) architecture.
  • Massive 1 million token context window.
  • Native vision capabilities.
  • Strong early performance in coding, agentic tasks, and long-context reasoning (competitive with or close to top closed models like GPT-5.6 variants or Claude models in certain benchmarks).
  • Pricing: $3 input / $15 output per million tokens (with caching discounts).
  • Important: Full weights are scheduled for open release (modified MIT license) around July 27, 2026 โ€” making it one of the first truly open frontier-scale models.

It represents a step up from prior Kimi models and adds meaningful competition from a non-U.S. lab.1

The Two Images in the Tweet

These are screenshots from Artificial Analysis (an independent benchmarking site):

  1. Bar chart โ€” โ€œCost per Intelligence Index Taskโ€ (lower is better). It breaks down costs by token type (input, output, reasoning, cache, etc.) and shows Kimi K3 as relatively expensive compared to some GPT-5.6 variants.
  2. Scatter plot โ€” Intelligence Index (y-axis) vs. Cost per task (x-axis, log scale). It positions models in quadrants, with the โ€œmost attractiveโ€ area being high intelligence at low cost. Kimi K3 appears competitive on intelligence but not the cheapest.

These visuals support Bakerโ€™s point that while Kimi K3 reaches frontier-level intelligence, it is not the most token-efficient, making its effective cost per useful output higher than more efficient models.

Bakerโ€™s Core Argument (Broken Down)

Bakerโ€™s thesis is economic and structural, not just technical:

  • A world with only 2โ€“3 dominant closed frontier labs (with ~90% inference margins) is bad for almost everyone else.
    Those labs would act as monopsonies (dominant buyers) for power, chips, data centers, and compute. They would vertically integrate upward (into infrastructure) and downward (into apps/software), capturing most of the value. Competition at the model layer prevents this concentration.
  • Anything that increases competition or lowers margins at the model layer is broadly positive.
    It redistributes economic value to the rest of the stack: more demand for GPUs/chips (inference especially), power, data centers, hyperscalers/neoclouds, and software (because cheaper/better AI drives wider adoption).
  • Token efficiency matters more than raw price per token.
    Baker notes Kimi K3 is roughly the same per-token price as GPT-5.6 but 50-70% more expensive overall according to Artificial Analysis. It is a โ€œtoken wastrelโ€ โ€” it uses more tokens to achieve similar results (lower โ€œintelligence density per tokenโ€).
    โ†’ True winners deliver the most intelligence per dollar. Cost-per-token + efficiency together determine real value.
  • Why NVIDIAโ€™s Jensen Huang loves open-source models.
    Open models still require massive compute to run at scale (same as closed ones). More models + wider adoption = more inference demand = more GPUs sold. A few ultra-high-margin closed labs would concentrate spend and reduce overall volume.
  • Vertical integration changes the game.
    Companies like Google, Meta, or xAI/SpaceX that have both strong models and their own infrastructure donโ€™t obsess over model-layer margins โ€” they capture value elsewhere. This is why more competition (open or vertically integrated) hurts pure-play high-margin model companies like OpenAI/Anthropic.
Why Itโ€™s Only โ€œPotentiallyโ€ Negative for OpenAI and Anthropic

Baker is careful:

  • Their products, user interfaces, and โ€œharnessesโ€ (the surrounding tools/ecosystems) may matter more than raw model performance right now.
  • They likely have more advanced internal models/checkpoints (possibly already used for recursive self-improvement/RSI).
  • Kimi K3โ€™s token inefficiency softens the blow for now. A more efficient open-source frontier model (or stronger showings from Grok, etc.) would be more threatening.
  • OpenAI and Anthropic are already moving to vertically integrate.
Overall Takeaway from the Post

Baker sees Kimi K3 (and the wave of strong Chinese models like Qwen, DeepSeek, GLM) as healthy competition that prevents unhealthy concentration. Itโ€™s not yet a full โ€œSputnik momentโ€ (a truly dominant, efficient open-source frontier model), but it points in that direction.

Net effect he describes: Lower % margins at the model layer โ†’ more total margin dollars flowing to infrastructure, chips, power, and software layers. This is bullish for the broader AI ecosystem even if it pressures the pure closed frontier labs.

The tone is analytical and balanced โ€” he acknowledges uncertainties and that time will tell (likely quickly). Replies to the post generally agree this kind of competition is positive for the industry overall and could shift more spend toward efficient inference hardware.

This fits Bakerโ€™s broader investing views on AI: he often emphasizes infrastructure winners, token economics, and how value accrues across the full stack rather than just at the smartest model.

Interesting.

Categories
AI

Memory Optimization

After thinking about the very high margins that Micron recently reported (greater than Nvidiaโ€™s) I asked Grok 4.5:

What startups are pursuing technology that reduces memory requirements for ai models and inference?

Several startups are actively developing technologies to slash memory requirements for AI models and inference. These efforts target the core bottlenecks in large language models (LLMs) and multimodal systemsโ€”primarily through extreme quantization (reducing bit precision of weights), alternative architectures that avoid quadratic memory scaling (e.g., state-space models or hybrids), sparsity/partial activation (like Mixture-of-Experts or neuron-level sparsity), efficient runtimes/kernels, and hardware innovations like in-memory computing that minimize data movement.

This space is especially active for on-device/edge AI (phones, laptops, embedded systems) and cost/energy-efficient cloud inference, where fitting larger or more capable models into limited RAM/VRAM, reducing KV cache bloat in Transformers, or lowering bandwidth demands delivers major gains.

Here are some of the most prominent startups and their approaches (as of mid-2026):

Extreme Quantization & Model Compression

  • PrismML (Caltech spinout, backed by Khosla Ventures): Focuses on extreme low-bit architecturesโ€”1-bit binary (โ€œBonsaiโ€) and ternary (1.58-bit, weights in {-1, 0, +1}). They compressed Alibabaโ€™s Qwen3.6 27B model from ~54 GB (FP16) to under 4 GB and ran the full dense 27B model on an iPhone 17 Pro. Claims include up to 14ร— smaller memory footprint, 8ร— faster inference, and significantly lower energy use, with competitive or better benchmark performance. They have open-sourced Bonsai models (including smaller 8B/4B/1.7B variants) under Apache 2.0 and are in discussions with Apple. This represents one of the most aggressive commercial pushes into 1-bit/ternary models for on-device deployment.
  • Mobius Labs (Berlin): Developed Half-Quadratic Quantization (HQQ), a fast, calibration-light post-training quantization method that enables high-accuracy low-bit models (including aggressive 2-4 bit). They demonstrated quantizing Llama 70B to run on a single GPU instead of four without major accuracy loss, directly cutting memory and compute needs. Their work extends to FP4 optimizations and integrates with frameworks like vLLM.
Alternative Architectures for Inherent Memory Efficiency
  • Liquid AI (MIT spinoff): Builds Liquid Foundation Models (LFM / LFM2 series)โ€”hybrid architectures combining gated short convolutions with grouped-query attention (GQA) blocks, plus MoE variants. These deliver substantially lower memory footprints than Transformers (especially for long contexts, avoiding massive KV cache growth), faster prefill/decode (up to 2ร— on CPU in some cases), and strong on-device performance. Examples include tiny models (230Mโ€“350M params, often
  • Cartesia: Specializes in state-space models (SSMs) and Mamba-derived/hybrid architectures. These offer near-linear (or better) scaling with sequence length and near-constant memory complexity during inference (fixed or minimal growth vs. quadratic KV cache in Transformers). This enables efficient long-context processing on smaller devices with lower energy use. They provide on-device libraries (e.g., โ€œEdgeโ€) and models like hybrid Mamba-2 variants optimized for throughput and fixed memory footprints.
On-Device/Edge Runtimes & Optimizations
  • Cactus (Y Combinator-backed): Builds cross-platform, energy-efficient inference runtimes and kernels for mobile, wearables, and low-power devices. Supports aggressive quantization (down to 2-bit), delivers very low latency (sub-50 ms time-to-first-token), and keeps models small (e.g., hundreds of MB for small models). Emphasizes privacy (fully local) and broad model support (Llama, Qwen, Gemma, etc.).
  • Embedl: Focuses on edge-optimized multimodal models and mixed-precision quantization (e.g., W4A16-Edge2 with selective FP16 layers + โ€œFlashHeadโ€ optimizations). Enables fast multimodal reasoning on devices with under 8 GB RAM (e.g., Jetson Orin Nano/AGX), with significant speedups over baselines while recovering most accuracy.
Hardware Innovations Reducing Memory Pressure
  • d-Matrix: Develops digital in-memory computing (DIMC / 3DIMC) chips and platforms (e.g., Corsair) that integrate compute directly with/near memory (initially SRAM, now advancing to 3D DRAM). This drastically cuts data movementโ€”the dominant memory bandwidth and energy bottleneck in inference. Claims include major speedups (up to 10ร— in some workloads) and better efficiency vs. traditional GPU/HBM setups, enabling higher performance with more manageable overall memory system requirements.
  • Hailo: Designs efficient edge AI inference chips (e.g., Hailo-10H series) optimized for low-power LLM and multimodal workloads. Their chips target 1โ€“3B parameter models at very low power envelopes (~2โ€“2.5 W), making deployment feasible on memory- and power-constrained edge devices where traditional hardware would struggle.
Other Notable Mentions
  • Tiiny AI: Uses neuron-level sparse activation (โ€œTurboSparseโ€) + heterogeneous engines for running very large models (up to 120B) fully on-device/offline in compact hardware.
  • Broader ecosystem players (inference platforms like Together AI or Fireworks) heavily optimize serving (quantization, caching, KV cache management) but are more infrastructure-focused than core technology innovators in compression/architecture.
Key Trends and Context

Quantization (especially post-training or quantization-aware) remains the most immediate lever and is widely adopted. Architectural shifts (SSMs, hybrids, MoE with low active parameters, convolutions) address fundamental scaling issues like KV cache memory growth. Hardware approaches (in-memory compute, high on-chip SRAM, efficient NPUs) tackle bandwidth and data movement. Many solutions combine these (e.g., architecture + quantization + sparsity).

Progress is rapid, driven by on-device AI demands (Appleโ€™s interest in PrismML is notable) and the economics of scaling inference. Many models are open-sourced or have open weights, accelerating adoption. Results vary by workloadโ€”extreme 1-bit models excel in certain scenarios but may need custom kernels; SSM/hybrid models shine for long context or linear scaling needs.

For the absolute latest benchmarks or specific model deployments, check company sites (e.g., PrismMLโ€™s Bonsai releases, Liquid AIโ€™s LFM models on Hugging Face) or recent evaluations, as the field evolves quickly.

Interesting!

Categories
AI Business

The Reverse Information Paradox We’ve Always Had

Satya Nadella wrote recently about what he calls the Reverse Information Paradox: enterprises pay for AI intelligence twice. Once in money. Again in the proprietary knowledge they surrender through every prompt, correction, and evaluation. The better they use the model, the more of their own institutional understanding leaks into someone else’s system. The vendor ends up knowing more about the buyer’s business than the buyer knows about what the vendor retained.

Replace “model” with “employee” (or โ€œconsultantโ€) and the paradox is not new at all.

You pay for a person once with salary. You pay again with something harder to price: the context, relationships, and judgment they must absorb to become useful to you. The better they perform, the deeper the immersion, the more of your particular way of doing things moves into their head. Every correction and late-night conversation is another trace of institutional memory changing hands. When they leave, some of that memory leaves with them. Not always through theft. Usually just through the ordinary residue of good work.

The visible cost is salary; the invisible cost is the slow transfer of what makes you distinctive. High performers get more access precisely because they’re high performers, which means the leakage accelerates exactly when you can least afford it. The exhaust is just harder to see with people than with tokens โ€” it moves through conversation and mental models instead of logs.

The analogy has a limit, and the limit matters. Employees bring knowledge in, not just absorb it. They have judgment and relationships a model doesn’t. Models are purely absorptive, and once something is inside them, it’s infinitely reproducible โ€” a person can only be in one place, working for one employer, at a time. We’ve had a few hundred years to build tools for the human version of this problem: contracts, culture, non-competes. The model equivalent is still being invented in real time, which is exactly why Nadella felt the need to name it.

Apple’s recent legal action against former employees who joined OpenAI is this pattern in its sharpest form. Whatever the specifics, the shape is familiar: people who spent years inside one of the most sophisticated organizations in the world, carrying out knowledge that never appeared on any balance sheet and was hard to contain. No one fully anticipates what a mind absorbs simply by being in the room long enough.

That’s the real difference between the silicon case and the human one. You can try to take action to wall off knowledge flowing to a model. You cannot wall off what someone has learned to notice.

Categories
AI Apple Google

The Library You Already Own

Sharon Park in the morning is not a dramatic place. There’s a duck pond, a stand of oaks that go gold too briefly in November, and a loop I’ve walked enough times that my legs know it better than my eyes do. It is, in other words, exactly the kind of place where a person starts talking to himself. Not out loud. In the productive, low-grade way โ€” turning a sentence over, arguing with an idea from the day before, checking a thought against something you believe about yourself.

I think in five years I’ll be doing that walk with something else along. Not a search engine. Not another chatbot trained to know a little about everything and a lot about nothing in particular. Something closer to a second set of eyes on my own life โ€” a reasoning engine, lean and mostly private, that has actually read the things I’ve written and doesn’t need me to explain who I am before it’s useful.

Here’s the distinction that matters, and it took me longer than it should have to see it clearly. The AI industry has spent years in an arms race over how much of the world a model can hold โ€” more facts, more languages, more of the internet compressed into weights. That race will keep going, and somebody else can have it. What I want is smaller and stranger: a model that knows comparatively little about the world and quite a lot about me. My core values document. The portfolio spreadsheets. Fifteen years of blog posts. The half-finished notes for the I-280 project, sitting in a folder, waiting for someone โ€” or something โ€” to ask the right question about them.

I spent a career in payments infrastructure, which means I spent a career thinking about a very specific kind of trust: the kind where a stranger’s system has to make a judgment call, in milliseconds, about whether to say yes. Fraud models don’t work because they know everything about commerce. They work because they know an enormous amount about one account, one pattern, one person’s ordinary Tuesday โ€” enough to notice when Tuesday stops being ordinary. That’s the architecture I keep picturing, aimed inward instead of outward. Not a system trying to know the world. A system trying to know me, well enough to notice when I’m drifting from what I said I cared about.

I can already feel the shape of the mornings this would change. Right now, when I sit down to look at RMD requirements against the tax picture, I’m doing the translation myself โ€” pulling numbers into a story I can actually feel the weight of. A reasoning engine grounded in my real holdings wouldn’t just run the scenario. It would know that I don’t want the scenario dressed up as a spreadsheet; I want it dressed up as a conversation, unhurried, the kind you’d have over lunch with someone who already knows the whole situation. And on the mornings when I sit down to write, instead of staring at a blinking cursor and a blank page that has no idea I exist, I’d be handing a draft to something that has actually read my last two hundred posts and knows the difference between the sentence I’d write and the sentence I’d cut.

None of this is especially exotic technology. Apple and Google are already building toward it โ€” Neural Engines fast enough to do real reasoning on-device, retrieval systems that can reach into your own files instead of the entire internet, fine-tuning that’s getting cheap enough to personalize rather than merely customize. The more interesting story here isn’t privacy, though privacy is real. It’s architectural: what happens when the expensive, impressive part of the system โ€” the part that knows everything โ€” becomes optional, and the cheap, personal part โ€” the part that knows you โ€” becomes the whole point.

What I don’t yet know is what this will cost me. A tool that reasons this well about my own life is also a tool I could lean on instead of doing the leaning myself, and there’s a version of this future where the walk around Sharon Park stops being mine and starts being a conversation with something that finishes my sentences a little too well. I’d want some way of knowing, plainly, what it’s drawing from and what it’s guessing at โ€” less a nutrition label than a kind of honesty I could check against, the way you’d check a fraud model’s confidence score before you trusted it with a yes.

But most mornings, I think I’d take the trade. Not because I want to think less. Because for thirty years I’ve been collecting the raw material โ€” the notebooks, the portfolios, the half-built essays โ€” and it would be something, finally, to walk beside a mind that had actually done the reading.

Categories
AI

The Taste Beneath the Summary

The real work of staying informed has never been volume. It has been the quiet, repeated acts of judgment: does this matter, to whom, why now, what is the signal beneath the noise.

A recent piece from Bridgewater’s AIA Labs and Thinking Machines Lab, “Learning to Replicate Expert Judgment in Financial Tasks,” describes training models to do the triage investors actually doโ€”filtering news, research, central bank documents, internal notes, for relevance. Frontier models struggled with judgments that looked simple and weren’t. The fix wasn’t a bigger model. It was Qwen, fine-tuned on labeled examples from practitioners, and it beat the frontier leaders while costing a fraction to run.

The bottleneck was never model size. It was taste. And taste, it turns out, can be taught to something small and cheap, if you’re precise enough about what you’re teaching itโ€”a market’s worth of Mercors is already proving the same thing at scale.

The researchers were clear that expert judgment doesn’t reduce to rules or prompts. It took high-quality, domain-specific labels from people doing the actual work. The most powerful systems will be built in partnership with practitioners who can say, and keep saying, what “good” looks like in their own context.

Which raises the question I haven’t answered yet: what would I actually put in the labels, if someone asked me to teach my own taste to a cheap model.

Categories
AI

The Quiet Setup: MacSparkyโ€™s Robot Assistant and the Unfair Advantage Still Available

A single X post caught my attention this week. It described something quietly happening among a small group of solo professionals. They arenโ€™t working longer hours or grinding harder. Instead, theyโ€™ve built a particular kind of setup around AI that carries much of the load.

While most of us still treat powerful models as clever search barsโ€”typing questions and copying answersโ€”these folks have given the AI a rich folder of context, a briefing file that orients it to their world, connections to their tools, and routines that let it produce real work on its own. The result can look like the output of a small team. From the outside it reads as talent or luck. Up close, itโ€™s mostly architecture.0

The post (from @zephyr_hg) emphasized that this advantage remains available because most people havenโ€™t yet made the shift from one-off prompting to building persistent systems. It landed with me because it echoes so closely the practical territory David Sparks (MacSparky) has been mapping for months in his Robot Assistant Field Guide.

MacSparkyโ€™s Approach: From Chatbot to Persistent Colleague

Davidโ€™s work centers on building a true personal assistant using Obsidian (for a local, plain-text knowledge base) and Claude (in its file-aware โ€œCoworkโ€ or project capabilities). The system isnโ€™t a chatbot that forgets everything between conversations. Itโ€™s designed to remember your projects, preferences, and people; triage email in your voice; handle morning briefings; track tasks; process documents; and support weekly reviewsโ€”freeing you from what David calls the โ€œdonkey work.โ€

The key ingredients will sound familiar to anyone who read that X post:

  • A dedicated context layer (your Obsidian vault or structured folder) holding the details of how you work.
  • Briefing/instruction files that tell the model who you are and what good looks like.
  • Integrations that connect it to email, calendar, files, and other tools.
  • Skills and routines that turn one-time intentions into repeatable, low-friction action.

David has been refreshingly transparent about the journey. He experimented earlier with more fully autonomous agents and even shut one down after learning what felt reliable and aligned. The Robot Assistant Field Guide distills those lessons into videos, workshops, templates, and a starter kit that lets people build without needing to code.

Why This Matters Now

Both perspectives point to the same shift in stance: moving from โ€œHow do I prompt better today?โ€ to โ€œWhat kind of system do I want running alongside me every day?โ€

For me, at this stage of life, that question carries weight. Iโ€™m not chasing maximum output for its own sake. I want arrangements that protect attention and energy for what actually mattersโ€”deep reflection, family history work, thoughtful investing, writing that might be useful to others, and simply being present. A well-designed AI setup doesnโ€™t just save minutes; it changes the texture of the day by reducing context-switching and repeated explanations.

It feels like finding a productive seam in the current moment of AI evolutionโ€”one of those hidden transitions where leverage quietly compounds if youโ€™re willing to build the architecture.

The Door Remains Open

The encouraging message in both the X post and Davidโ€™s teaching is that this isnโ€™t locked behind rare talent or expensive infrastructure. The models are accessible. The patterns are becoming clearer. Whatโ€™s required is the decision to treat AI less like a toy and more like a colleague youโ€™re willing to orient and trust with real work.

I donโ€™t have my own โ€œrobot assistantโ€ fully built yet. Iโ€™ve been experimenting with custom agents, structured daily scans, and ideas like โ€œThe Observatoryโ€ for signal synthesis. Reading these sources side-by-side sharpened my sense of the next layer: giving the system a proper home, clear instructions, and meaningful recurring work.

If youโ€™re a solo professional, creator, or lifelong learner feeling the press of too many small tasks, this is worth exploring. Start small. Build a modest context folder. Write a briefing file that captures how you think. Experiment with one routine. Iterate from there.

The setup that outworks the grind isnโ€™t magic. Itโ€™s deliberate, learnable, and still wide open.


What setups are you experimenting with these days? Iโ€™d love to hear in the comments or on X.


Categories
AI Podcasts

A Remarkable Conversationโ€ฆ

Highly recommend this conversation between Harry Stebbings and Clay Bavor. Among many topics, I especially enjoyed the discussion about not investing in frontier models, the important values, the particular importance of craftsmanship, intensity, and family. And the special conversation about parenting and kids near the end. Just a delightful conversation to be able to enjoy!

Key Highlights:

โ€ข Founding Sierra: Bavor explains why he and Taylor chose to start Sierra, focusing on the transformative potential of language model-based agents (1:37 – 5:53).
โ€ข The AI Tech Stack: Sierra focuses on building enterprise-grade agent architectures and fine-tuning models on top of open-weights models rather than pre-training foundation models from scratch, prioritizing capital efficiency (5:53 – 7:15).
โ€ข Unbounded Demand for Intelligence: Bavor argues that there is massive, unmet demand for “frontier-level” intelligence in fields like coding, science, and legal work (7:15 – 11:41).
โ€ข Internal AI Operations: He details the use of Pinecone, an internal AI agent Sierra developed to navigate company data, streamline engineering, and assist in recruitment (18:36 – 22:00).
โ€ข Enterprise Strategy: Sierra employs a “forward-deployed” engineering model, embedding staff within client companies to ensure rapid, effective integration of AI, leading to quick deployment timelines (30:12 – 33:22).
โ€ข Board Governance: To keep pace with the speed of AI development, Sierra operates on a six-week board meeting cadence, utilizing comprehensive memos instead of traditional slide decks (39:07 – 41:13).
โ€ข Corporate Culture: Bavor emphasizes values like craftsmanship, intensity, and family. He also highlights the importance of working in-person to foster apprenticeship, mentorship, and a cohesive team culture (43:02 – 55:41).

Categories
AI Apple Google

The Floor

I compared the frontier to a three-star chef making grilled cheese in “Context Rot” โ€” the smartest models on earth spending most of their time on work beneath them, the way a chef trained at Le Bernardin might still melt cheese between two slices of bread on a Tuesday night and call it dinner. The comfort was the point: if the sharpest tool is saved for hard problems and something merely-very-good handles the rest, nobody’s losing anything. The floor was never the interesting part.

I’ve kept turning the joke over, and I think I had the wrong worry.

Watch what companies do with their AI spend, not what they say. Coinbase moved engineers off frontier models onto open weights and cut its AI spend nearly in half while usage kept climbing. Nvidia runs a closed model as orchestrator and routes the actual volume โ€” the daily uncelebrated bulk of it โ€” to open weights it controls. The frontier is becoming a dispatcher, deciding where the request goes and rarely doing the work itself. The instinct is to worry about whose open weights end up running that volume, and right now the most capable ones at scale are Chinese โ€” GLM, Kimi โ€” which makes it tempting to read this as a contest America is quietly losing: the floor of the AI economy built somewhere else, at a price export controls can’t touch. You cannot embargo a file already downloaded. You cannot price-match free.

But that framing has a hole. Google’s own Gemma family is open-weight and good enough to handle that daily volume without anyone reaching for GLM or Kimi. “Open weights are a Chinese story” only holds if you don’t count the open models the company running Android and half the internet’s search traffic has already shipped.

And once I saw that hole, a bigger one opened behind it. I’ve been trying Apple’s new Siri โ€” arriving with iOS 27 this fall, genuinely surprisingly good in beta โ€” and it made me realize open weights, of any nationality, were never going to cook most of the world’s dinners. Apple and Google are.

Consider what actually determines where the world’s routine inference runs. Not which model benchmarks best, not which weights are downloadable โ€” what’s already installed. Apple ships to well over a billion active devices before routing a single query through Siri’s new architecture. Nobody has to be persuaded to try it, or hear about it on a podcast; it’s the thing that answers when you press the button you’ve pressed for a decade. Google owns the search bar and the Android default the same way. Between them, that’s most of the world’s phones โ€” and phones are where most of the world’s questions get asked.

The open-weight framing assumes the floor is up for grabs, that whoever ships the best free model wins the daily grind by merit. But the floor was never a bazaar. It’s a set of defaults, owned by whoever already has the device in your hand, not whoever holds the most generous license. Apple didn’t need to win the model war to win this. Its heaviest reasoning tier is built with Google, running on Nvidia chips in Google’s cloud, under a deal reported at roughly a billion dollars a year โ€” Apple doesn’t fully own the engine doing the thinking. It doesn’t need to. It owns the button.

That’s a quieter concentration than an export-controls fight, and a harder one to dislodge. An open model can be forked, distilled, undercut, or out-competed by the next release. A billion phones with an assistant built into the lock screen cannot be routed around. Whoever’s weights hum underneath barely matters, the way it barely matters to a diner which supplier delivered the flour. What matters is whose kitchen the meal came from, and whose name is on the door.

The grilled-cheese chef was never the risk. Two chefs are about to own nearly every kitchen on earth, and most of us will never notice โ€” because a kitchen you’ve been eating out of for a decade doesn’t feel like something that was won. It just feels like home.

Owning the kitchen and getting paid for what’s cooked in it, though, turn out to be two different questions. That one’s for another post.

Categories
AI AI: Inference Semiconductors Uncategorized

5 Critical Management Lessons from the Founders at Etched

How two young founders are building what could become one of the most important companies in the AI era โ€” and what their story teaches about leadership, execution, and building at the edge of the possible.

I recently listened to the latest Invest Like the Best podcast from Patrick O’Shaughnessey which was a remarkable conversation with Gavin and Rob, the founders of Etched, the company building specialized AI inference hardware that’s aiming to be radically better than existing solutions. Their story โ€” starting as very young founders against massive skepticism, raising serious capital, and now shipping full rack-scale systems โ€” is packed with hard-earned wisdom.

One of the comments Patrick makes at the beginning was how during his due diligence on the company he kept being told that semiconductor technology wasn’t a place for young people. You need seasoned, middle age experts to master this domain. Exactly not these founders.

Note: the following is based upon an AI’s analysis of the conversation transcript with me asking “What are the five most important management lessons from this conversation?” These lessons are relevant whether you’re leading a team, building a product, or simply trying to do meaningful work in our fast-moving world.

1. Velocity Compounds โ€” Prioritize Speed Ruthlessly

In hardware, and increasingly in any deep-tech endeavor, speed isn’t just an advantage; it’s often the deciding factor.

Etched didn’t just design a chip โ€” they built the full inference solution (chip, board, power delivery, interconnects, cold plates, and production processes) in parallel. They sent engineers to live in Bangalore for months to unblock vendors. They ran 24/7 shifts and did massive pre-work (including putting full chip designs on FPGA clusters) so that when the silicon finally arrived, they had working inference in racks in just 40 days.

Key takeaway: Look for every opportunity to parallelize. Accept higher short-term costs if they buy meaningful time. As they put it, “You win by shipping.” The best part is often no part โ€” and the best vendor is no vendor, when vertical integration lets you move faster. Velocity, velocity, velocity.

2. Build Teams with Legends + High-Drive Talent

One of the most distinctive parts of their approach is how they recruit. They seek out “Legends” โ€” people who have done the hardest versions of the problem before (like the engineer who built Nvidia’s HGX and DGX systems) โ€” and pair them with exceptionally driven, somewhat naive high-performers who refuse to accept conventional limits.

They use “project-based recruiting,” mapping the hardest technical problems ever solved and persistently pursuing the actual people who did the real work. Their culture self-selects for people willing to move their families to San Jose to bet on two young founders taking on the world.

Key takeaway: For breakthrough work, average talent doesn’t suffice. The combination of deep experience and raw, first-principles energy creates magic. Invest heavily in finding and retaining these people โ€” even if it takes 20 conversations. You can also learn a lot if the best in the world talent turns down the opportunity to work with you!

3. Assume It’s Possible, Then Solve the “Unsolvable” Problems

Repeatedly in their story, experts told them certain things were impossible. Their response? Assume it is possible and figure out how.

The most striking example was a clock domain crossing issue that required aligning signals to within 50 picoseconds โ€” something many engineers said couldn’t be done. People quit. They solved it in about two weeks during a very dark period.

Key takeaway: When you hear “impossible,” treat it as the beginning of the investigation, not the end. Cultivate a “find a way” mindset across the team. The moments when things feel hopeless are often when the most important progress happens. I’m constantly struck by how often persistence results from simply realizing (or assuming) that something is actually possible.

4. Production Is the Real Product

Etched’s mantra is “Production is the product.” They obsess over not just technical performance but manufacturability, supply chain resilience, serviceability, and the ability to scale to gigawatts.

They made deliberate choices around process nodes and memory to avoid zero-sum competition. They built their own factory processes and test infrastructure early. Future designs are being simplified specifically for faster production cycles and higher reliability at massive scale.

Key takeaway: In any business that hopes to reach real scale, think end-to-end from the beginning. Technical excellence without production excellence is just a prototype. Optimize for output (tokens, units, whatever your metric is) at volume. There’s a lot of “zero to one” thinking here.

5. Bet Big and Stay Existentially Focused

Building in semiconductors requires enormous capital. Etched raised roughly $100 million early on when they were still very young and pre-tapeout โ€” after most traditional investors had passed. They knew half-measures wouldn’t work.

This existential focus (this one product determines whether the company lives or dies) creates a different level of intensity that attracts talent, suppliers, and customers who believe.

Key takeaway: Match your ambition with appropriate resources and commitment. Clear existential stakes help filter for the right people and partners. In a world of distractions, singular focus on what truly matters is a superpower.

Final Thoughts

Gavin and Rob’s story is the combination of technical sophistication and deep human resilience. They faced a tough personal battle with cancer (in Rob’s case), widespread doubt, brutal technical challenges, and fundraising pressure โ€” and kept moving forward with curiosity, determination, and humility.

In an age of AI and accelerating technology, the ability to build teams that can solve seemingly impossible problems at speed may be one of the most valuable capabilities a leader can develop. Their example reminds us that the future belongs not just to the smartest, but to those who can execute with urgency while maintaining clear principles. Velocity, velocity, velocity.