The TWIML AI Podcast

Diogo Almeida on why you get the AI you optimize for

Diogo Almeida· Co-founder & CEO of TypeSafe AI at TypeSafe AI
·~90 min·English·TWIML
LLMTrainingInferenceAI Company
TL;DR

TypeSafe's Diogo Almeida argues AI is useless for everyday automation because we optimized it to write strings for humans, and that the fix is a new training objective — calibrated decisions software can trust like SQL.

01Core thesis

You get the AI you optimize for

The gap between genius AI and useless AI is not intelligence — language models were optimized to produce strings for humans, while automation needs decisions a computer can act on.

you get what you optimize for and basically all the string LLMs have been optimized for strings and strings are meant to be consumed by humans or other LLMs. But if you want something like accounting that's meant to be consumed by a computer, right?

— Diogo Almeida, The TWIML AI Podcast
Key Insight
The framing quietly rejects the whole "AI is almost smart enough" narrative: the same model that aces research can't change a credit card not because it lacks IQ, but because nobody trained it toward the shape software needs.

02Reliability

Jaggedness is a choice, not a law of AI

A wrong refund and keyboard-mashing are equally wrong by accuracy, yet models never mash the keyboard — so the jagged edge comes from where training pressure was aimed, not from AI itself.

saying to give someone a refund when you shouldn't is from an accuracy perspective exactly the same as like mashing your face on the keyboard. You know it's it's the same 0% accuracy but you never get them mashing your face on the keyboard because of the asymmetry in the optimization.

— Diogo Almeida, The TWIML AI Podcast
Key Insight
If unreliability is manufactured by the optimization process rather than inherent, then it can be engineered back out — which is the entire premise for building a different kind of model instead of waiting for the next frontier LLM.

03ML worldview

The bitter lesson, re-pointed at the task

Almeida accepts that scale beats clever algorithms, but adds a twist: data matters more than compute, and choosing the right task to optimize matters most of all.

I believe that data matters a lot more than compute. So you actually need data to put the compute on and more important than data is you need the right task. So this is the most important thing in all of the all of the ML that we do

— Diogo Almeida, The TWIML AI Podcast
Key Insight
Pre-training loss was once a wild bet that paid off; naming a genuinely new task — not tuning the same one — is what produces the rare giant jumps in ML, which is how he positions machine-native intelligence against incremental benchmark chasing.

04The method

RLCD optimizes for decisions, not human preference

RLHF taught models to chase human preference and produce pleasing strings; his class of algorithms, RLCD, optimizes instead for calibrated decisions that software can depend on.

So like when I talk about RLCD, I talk about like the general class of algorithms that optimize specifically for calibrated decisions because that is what's useful for software.

— Diogo Almeida, The TWIML AI Podcast
Key Insight
He is careful that RLCD names a family of objectives, not one secret algorithm — inviting others to build in the same direction, the same way publishing RLHF let the whole field move once the task was named.

05Calibration

The probability is the product

Because a decision model exposes calibrated probabilities, the engineer — not the model — sets a tunable threshold for when to act, turning intelligence into a controllable software knob.

That is insane behavior from like from an engineering point of view, right? Why wouldn't you have a threshold? Why wouldn't you have a tunable threshold?

— Diogo Almeida, The TWIML AI Podcast
Key Insight
RLHF and RLVR trade calibration away for the overconfidence that fluent text needs, so a model that stays calibrated is doing something the mainstream pipeline actively destroys — and that calibration, not raw accuracy, is what lets software tune behavior per business.

06Architecture

Assembled from open weights, not pretrained from scratch

Jev skips doing its own pre-training — which he calls a bad deal for sublinear gains — and is instead stitched from several already-pretrained open-weight models, each carrying a slice of the internet's compressed intelligence.

I have described it as like a a Frankenstein's monster of models. Um, which I I haven't read the book, but I've been told is at least innocent, if not the good guy.

— Diogo Almeida, The TWIML AI Podcast
Key Insight
Declining to pre-train is both a budget admission and a thesis: the frontier already compressed the internet into reusable intelligence, so the leverage now is in re-shaping that core toward a new objective, not re-paying the pre-training bill.

07The north star

Reliability so boring it is like SQL

The goal is intelligence you trust without re-checking — a primitive as predictable as a SQL query or a logic gate, which is when software engineers get superpowers.

I want to be a paragon of making AI so reliable that it's boring like SQL, you know, like I want AI to be so predictable that you can like write queries without having to even run them against like eval sets

— Diogo Almeida, The TWIML AI Podcast
Key Insight
"Boring" is the ambition, not a hedge: most AI products sell excitement and demos, but a dependency you embed and forget has to be dull and predictable — reliability, not novelty, is what developers are actually paying for.

08Agents

The harness is a horseless carriage

Almeida rejects the agent "harness" as a human-shaped crutch; the model is a low-level primitive you drop into ordinary code, where rigid rails are a feature for automation meant to run forever.

a harness itself is just it's a very um um horseless carriage type thing of trying to turn the intelligence into something that looks like a human.

— Diogo Almeida, The TWIML AI Podcast
Key Insight
It is a contrarian bet against the agent-framework wave: Almeida's guess — and he stresses he lacks the labs' data — is that the recent jump in coding agents came from putting the task in the model's training distribution, not from the loop around it, so the durable value would lie in the intelligence and the code, not the harness.