Diogo Almeida on why you get the AI you optimize for
TypeSafe's Diogo Almeida argues AI is useless for everyday automation because we optimized it to write strings for humans, and that the fix is a new training objective — calibrated decisions software can trust like SQL.
You get the AI you optimize for
The gap between genius AI and useless AI is not intelligence — language models were optimized to produce strings for humans, while automation needs decisions a computer can act on.
you get what you optimize for and basically all the string LLMs have been optimized for strings and strings are meant to be consumed by humans or other LLMs. But if you want something like accounting that's meant to be consumed by a computer, right?
Jaggedness is a choice, not a law of AI
A wrong refund and keyboard-mashing are equally wrong by accuracy, yet models never mash the keyboard — so the jagged edge comes from where training pressure was aimed, not from AI itself.
saying to give someone a refund when you shouldn't is from an accuracy perspective exactly the same as like mashing your face on the keyboard. You know it's it's the same 0% accuracy but you never get them mashing your face on the keyboard because of the asymmetry in the optimization.
The bitter lesson, re-pointed at the task
Almeida accepts that scale beats clever algorithms, but adds a twist: data matters more than compute, and choosing the right task to optimize matters most of all.
I believe that data matters a lot more than compute. So you actually need data to put the compute on and more important than data is you need the right task. So this is the most important thing in all of the all of the ML that we do
RLCD optimizes for decisions, not human preference
RLHF taught models to chase human preference and produce pleasing strings; his class of algorithms, RLCD, optimizes instead for calibrated decisions that software can depend on.
So like when I talk about RLCD, I talk about like the general class of algorithms that optimize specifically for calibrated decisions because that is what's useful for software.
The probability is the product
Because a decision model exposes calibrated probabilities, the engineer — not the model — sets a tunable threshold for when to act, turning intelligence into a controllable software knob.
That is insane behavior from like from an engineering point of view, right? Why wouldn't you have a threshold? Why wouldn't you have a tunable threshold?
Assembled from open weights, not pretrained from scratch
Jev skips doing its own pre-training — which he calls a bad deal for sublinear gains — and is instead stitched from several already-pretrained open-weight models, each carrying a slice of the internet's compressed intelligence.
I have described it as like a a Frankenstein's monster of models. Um, which I I haven't read the book, but I've been told is at least innocent, if not the good guy.
Reliability so boring it is like SQL
The goal is intelligence you trust without re-checking — a primitive as predictable as a SQL query or a logic gate, which is when software engineers get superpowers.
I want to be a paragon of making AI so reliable that it's boring like SQL, you know, like I want AI to be so predictable that you can like write queries without having to even run them against like eval sets
The harness is a horseless carriage
Almeida rejects the agent "harness" as a human-shaped crutch; the model is a low-level primitive you drop into ordinary code, where rigid rails are a feature for automation meant to run forever.
a harness itself is just it's a very um um horseless carriage type thing of trying to turn the intelligence into something that looks like a human.