Balance of Power

Francis deSouza on evidence over hopes and fears

Francis deSouza· CEO of Scale AI at Scale AI
·~7 min·English·Bloomberg
AI SafetyPolicyOpen SourceTraining
TL;DR

The CEO of Scale AI argues that AI policy should be set by independently measuring what every model — frontier and open-weight — can actually do and where it is risky, rather than by either acceleration hopes or catastrophe fears.

01The Core Thesis

Evidence, Not Hopes and Fears

deSouza wants the pace of AI and the rules around it set by measured capability and risk, not by either optimism or dread.

instead of basing our strategy, our regulatory strategy, or even the pace at which you develop models on hopes and fears, we really need to get a clear understanding of what the models are capable of, what the risks are

— Francis deSouza, Balance of Power
Key Insight
The framing quietly reframes the whole fight: it treats how fast and under what rules as outputs of measurement. Whoever owns the measurement owns the debate — which is the position Scale AI is selling.

02The Debate He's Entering

Two Camps, Both Valid, Truth in Between

One side says move fast because AI can cure disease; the other says slow down because the risks are real; deSouza refuses to pick a pole and puts the answer in the middle.

We're seeing very smart people have very different opinions on this topic, and they're both valid, to be honest.

— Francis deSouza, Balance of Power
Key Insight
Calling both camps valid strips the moral high ground from each side and reframes the fight as a question of evidence rather than conviction — a framing that happens to make an independent testing layer, like Scale's, the natural place to settle it.

03The Method

You Can't Mitigate What You Can't See

deSouza's point here is a dependency, not a preference: you cannot choose a mitigation for a risk you have not first measured.

We can't mitigate something if we don't fully know what the risk is.

— Francis deSouza, Balance of Power
Key Insight
“Moving in the dark” reframes the timeline: if a model ships before its risks are mapped, every mitigation can only trail the capability — a reaction to a risk already released, never a precondition of release.

04The Blind Spot

Test Every Model, Not Just the Frontier

deSouza's sharpest point is that the safety conversation fixates on frontier models while a lot of people are actually running open-weight models that get far less scrutiny in the debate.

a lot of people are using open weights models, and we need to understand the capabilities of those models and the risks that those models pose

— Francis deSouza, Balance of Power
Key Insight
Open weights are the hardest part of the gap to walk back: once a model's weights are public they cannot be un-released, so a frontier-only debate can set rules for the next model while the open ones already in circulation stay untested.

05The Rigor

A Right Answer Isn't a Good Answer

deSouza shows a model can reach the correct diagnosis by skipping steps and guessing, so scoring only the answer misses whether the reasoning was sound.

just evaluating models and the accuracy of a diagnosis isn't enough, because some models can get there by skipping a bunch of steps and guessing the answer

— Francis deSouza, Balance of Power
Key Insight
A lucky guess and careful reasoning can post the identical accuracy score, so an outcome-only test cannot tell them apart. That is why deSouza evaluates the path to the answer — did the model follow the steps? — not just the answer itself.

06The Business Underneath

Selling the Reliability Layer

Everything deSouza argues for maps onto Scale's three lines of business — supplying data, wrapping guardrails around agentic workflows, and running evaluations — which he frames as one reliability layer under all AI.

we are about providing the reliability layer for AI and whether the governments do it or companies do it

— Francis deSouza, Balance of Power
Key Insight
Bundling data, guardrails, and evaluations into one “layer” turns three services into infrastructure — something you depend on continuously rather than audit once. It is also how deSouza argues demand only grows: more models, and more capable ones, mean more to measure.