Francis deSouza on evidence over hopes and fears
The CEO of Scale AI argues that AI policy should be set by independently measuring what every model — frontier and open-weight — can actually do and where it is risky, rather than by either acceleration hopes or catastrophe fears.
Evidence, Not Hopes and Fears
deSouza wants the pace of AI and the rules around it set by measured capability and risk, not by either optimism or dread.
instead of basing our strategy, our regulatory strategy, or even the pace at which you develop models on hopes and fears, we really need to get a clear understanding of what the models are capable of, what the risks are
Two Camps, Both Valid, Truth in Between
One side says move fast because AI can cure disease; the other says slow down because the risks are real; deSouza refuses to pick a pole and puts the answer in the middle.
We're seeing very smart people have very different opinions on this topic, and they're both valid, to be honest.
You Can't Mitigate What You Can't See
deSouza's point here is a dependency, not a preference: you cannot choose a mitigation for a risk you have not first measured.
We can't mitigate something if we don't fully know what the risk is.
Test Every Model, Not Just the Frontier
deSouza's sharpest point is that the safety conversation fixates on frontier models while a lot of people are actually running open-weight models that get far less scrutiny in the debate.
a lot of people are using open weights models, and we need to understand the capabilities of those models and the risks that those models pose
A Right Answer Isn't a Good Answer
deSouza shows a model can reach the correct diagnosis by skipping steps and guessing, so scoring only the answer misses whether the reasoning was sound.
just evaluating models and the accuracy of a diagnosis isn't enough, because some models can get there by skipping a bunch of steps and guessing the answer
Selling the Reliability Layer
Everything deSouza argues for maps onto Scale's three lines of business — supplying data, wrapping guardrails around agentic workflows, and running evaluations — which he frames as one reliability layer under all AI.
we are about providing the reliability layer for AI and whether the governments do it or companies do it