Categories
AI

The Taste Beneath the Summary

A small, cheap model fine-tuned on expert labels just outperformed the frontier leaders at the one thing that’s always been hardest to automate — taste.

The real work of staying informed has never been volume. It has been the quiet, repeated acts of judgment: does this matter, to whom, why now, what is the signal beneath the noise.

A recent piece from Bridgewater’s AIA Labs and Thinking Machines Lab, “Learning to Replicate Expert Judgment in Financial Tasks,” describes training models to do the triage investors actually do—filtering news, research, central bank documents, internal notes, for relevance. Frontier models struggled with judgments that looked simple and weren’t. The fix wasn’t a bigger model. It was Qwen, fine-tuned on labeled examples from practitioners, and it beat the frontier leaders while costing a fraction to run.

The bottleneck was never model size. It was taste. And taste, it turns out, can be taught to something small and cheap, if you’re precise enough about what you’re teaching it—a market’s worth of Mercors is already proving the same thing at scale.

The researchers were clear that expert judgment doesn’t reduce to rules or prompts. It took high-quality, domain-specific labels from people doing the actual work. The most powerful systems will be built in partnership with practitioners who can say, and keep saying, what “good” looks like in their own context.

Which raises the question I haven’t answered yet: what would I actually put in the labels, if someone asked me to teach my own taste to a cheap model.

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from Scott Loftesness

Subscribe now to keep reading and get access to the full archive.

Continue reading