Squawk Pod

Bret Taylor on why you'll stop paying per token

Bret Taylor· Chairman of OpenAI; Co-founder of Sierra at Sierra
·~54 min·English·CNBC
InferenceBusiness StrategyAgentsAI Company
TL;DR

OpenAI chairman and Sierra CEO Bret Taylor argues that as frontier intelligence becomes a rented commodity, enterprises should stop paying per token and start paying for outcomes — and build their moat on customer relationships, not models.

01The Frame

The Two Questions Every CEO Is Asking

Taylor talks to about a hundred CEOs a month, and <strong>two anxieties dominate: am I getting value for my token spend, and where is my moat when everyone can rent the same intelligence?</strong>

how can I know that I'm getting the value from this spend on tokens. And then I'll use the word sovereignty, which is just making sure like what is my competitive moat in the age of AI as well.

Bret Taylor, Squawk Pod
Key Insight
The two questions collapse into one. When the model itself is a commodity you rent, the only levers a non-model company still controls are cost-per-outcome and ownership of the customer relationship — everything else is bought off the shelf.

02Mental Model

A Token Is to Intelligence Like a Watt Is to Electricity

The tidy analogy — a token is a unit of intelligence, like a watt is a unit of power — makes 'token efficiency' easy to grasp, but <strong>Taylor breaks his own metaphor: not every token is equal.</strong>

a token is to intelligence like a watt is to electricity like a unit of intelligence. And token efficient means how many tokens does it take to complete a task

Bret Taylor, Squawk Pod
Key Insight
The crack in the analogy is the whole point. If one model's token buys better output than another's, then 'cheapest per token' and 'cheapest per finished task' stop agreeing — and only the second number ever shows up on the invoice.

03Myth vs Reality

Open Weights Don't Mean Cheaper

The 'open-weight models are cheaper' claim quietly swaps two different costs: <strong>open weights may lower the cost to train, but you still burn just as many tokens to run them — and frontier models are more token-efficient.</strong>

just having open weights isn't actually the main thing driving many of those costs.

Bret Taylor, Squawk Pod
Key Insight
Open weights change who is allowed to train a model, not what it costs to serve one. The bill that repeats every request is inference, and that is set by token efficiency — which is exactly where Taylor argues the frontier labs still lead.

04The Big Shift

Stop Paying Per Token. Pay Per Outcome.

Paying per token, Taylor says, is like paying for Gmail by the CPU cycle — <strong>the market is moving to paying for outcomes: a closed loan, an authorized procedure, not raw compute.</strong>

where I believe where the world is going is paying for outcomes.

Bret Taylor, Squawk Pod
Key Insight
The slogan only holds if the vendor can deliver the outcome for less than its price. Taylor's real bet is that an applied-AI company like Sierra can amortize the R&D for specialized low-latency models across every customer — something no single enterprise could justify alone.

05Model Routing

Driving a Ferrari to the Grocery Store

Most teams reach for the top model on every task — <strong>like driving a Ferrari to the grocery store — because routing the easy jobs to a cheaper model isn't automatic yet.</strong>

do you want to use the Ferrari or do you want to drive in the Honda today and most people are choosing the Ferrari

Bret Taylor, Squawk Pod
Key Insight
Taylor is careful that the model makers aren't 'forcing' overspend — the waste comes from a missing routing layer. Whoever builds automatic price-for-performance routing captures the savings the buyer is currently leaving on the table.

06Where We Are

The Applied-AI Market Is Stuck in 1997

We're in the build-it-yourself era of AI, Taylor says — <strong>like 1997, when a login website cost tens of millions; the fix is applied-AI companies that hide the tokens entirely.</strong>

if you go back to 97 people were spending tens of millions of dollars just to make a website where people could log in and now it's easy

Bret Taylor, Squawk Pod
Key Insight
The 1997 comparison is a business thesis in disguise. If applied-AI companies abstract the tokens away, the durable value migrates to those abstraction layers — Sierra, Harvey — and away from the enterprises still buying tokens by the barrel.

07The Durable Moat

Your Moat Is the Customer, Not the Model

When everyone can rent frontier intelligence, the model is a commodity — <strong>the durable moat is the compounding data from your own customer relationships, which no competitor gets to reuse.</strong>

intelligence is actually going to be quite democratized and you need to think about what are the things around it like your customer relationships that can be durable

Bret Taylor, Squawk Pod
Key Insight
'What's mine is mine' is the enterprise version of this moat — Taylor ties it to contractual guarantees that your prompts don't train someone else's model. So the moat is only as durable as that guarantee: own the relationship, but also own the data rights, or the moat leaks straight back to the frontier.