a16z Podcast

Vladimir Keil on why the end-to-end job beats the system of record

Vladimir Keil· Co-founder & CEO at Leo
·~59 min·English·a16z
AgentsAI CompanyBusiness Strategy
TL;DR

Using his procurement startup Leo as the case study, Vladimir Keil argues that a vertical AI company beats model-equipped incumbents by owning the end-to-end job the system of record never captures, climbing a ladder of judgment as customer trust compounds.

01The Stakes

One Missed Email, Hundreds of Millions Gone

<strong>The agent's job is to decide, not retrieve:</strong> one of 500 routine emails says a part will slip two weeks, and missing it can cost a multi-million-dollar project hundreds of millions, so the agent has to judge impact rather than just file the new date.

And if they missed this email, hundreds of millions of damage.

— Vladimir Keil, a16z Podcast
Key Insight
The agent's real job is triage: deciding which of 500 look-alike emails is the one carrying nine-figure consequences. That reframes procurement software from a filing cabinet into a risk engine, and it is the first task a pure retrieval tool cannot do.

02Core Thesis

The Iceberg Under '8K for Aluminum'

<strong>The system of record shows the tip; the real work is the iceberg beneath it:</strong> a single line like '8K for aluminum' hides 30 meetings, 500 emails and weeks of cost modeling, and most of procurement happens outside the ERP entirely.

the legacy incumbent is limited to their system of record

— Seema Amble, a16z Podcast
Key Insight
If most of the work lives outside the system of record, then the incumbent's greatest asset (owning that record) is also its ceiling. The startup's data moat is the exhaust of doing the whole job end to end, not a database it already bought.

03Mental Model

Four Agents, One Ladder of Judgment

<strong>Agents climb a ladder of judgment, retrieval to process to policy to principal:</strong> incumbents slapped a chatbot on the record and stalled on the bottom rungs, while the value sits on the higher rungs, where the software must exercise judgment a strict rule cannot cover.

They're very much still limited to, I would say, retrieval and a little bit of process. They've not gotten into more judgment.

— Seema Amble, a16z Podcast
Key Insight
The ladder is a trust gradient dressed up as a capability gradient. Each higher rung means the software makes a judgment call the customer could later blame the vendor for, and that disincentive compounds with the internal product conflicts that keep incumbents on the bottom rungs.

04Why There Is Room

Held Back, Not Holding Back

<strong>Incumbents are not holding back, they are held back:</strong> selling a workflow tool and resolving the work are two different products for two different buyers, so the same installed base that makes selling easy also blocks the end-to-end build.

they're holding back but they are held back.

— Seema Amble, a16z Podcast
Key Insight
The incumbent's distribution advantage and its org chart are the same fact seen from two sides: the customer base that makes cross-selling trivial is also the political map that makes one team building a product that eats another team's product impossible.

05Earning Trust

Trust Is Earned One Negotiation at a Time

<strong>Autonomy is earned, not switched on:</strong> no enterprise starts with a fully autonomous negotiation agent, so a human stays in the loop feeding the agent feedback until trust compounds across 10k, 20k, then 100k negotiations.

No company and no enterprise starts with fully autonomous negotiation agents from day one.

— Vladimir Keil, a16z Podcast
Key Insight
Human-in-the-loop is not a safety compromise on the road to autonomy; it is the mechanism that produces autonomy. Each supervised negotiation is a labeled example tuned to how one specific enterprise operates, which a generic chatbot never collects.

06The Moat

The 70% Trap

<strong>A fast build is not production automation:</strong> Keil says engineers can rebuild Leo's first use case in about eight hours yet still reach only 70%, and 70% performance is not 70% automation, because the last stretch of exceptions, integrations and vertical data is most of the effort.

they can get to 80%, but those last 20% really matter.

— Vladimir Keil, a16z Podcast
Key Insight
70% performance is a vanity metric: if a human must still review everything, the enterprise has bought a second task, not automation. Seema's Fortune 500 cash-collection example shows a related barrier: too little quality context, plus the cost of maintaining mappings across changing ERPs, made the internal build not worth it.

07Net-New Value

Negotiating What Nobody Negotiated

<strong>The biggest wins hide in spend no one bothered to negotiate:</strong> enterprises often left anything under 50k unnegotiated for lack of capacity, so an agent can capture it with little downside, though relationship-sensitive vendors still need a human.

what is the risk now of having a bad negotiation agent? Nearly zero

— Vladimir Keil, a16z Podcast
Key Insight
The un-negotiated spend is nearly free money precisely because no human was ever assigned to it, so there is no job being displaced and little downside to a merely-okay agent, as long as the relationship itself is not the point. The cheapest wins are the ones no incumbent ever bothered to measure.

08Why Procurement

Boring, Emotional, and Enormous

<strong>Boring, emotional and enormous is the wedge:</strong> procurement is universally disliked and has barely changed in 20 years, yet 1% of savings is worth roughly 10% more sales, which is why Keil calls it a trillion-dollar opening.

boring, highly emotional and then plus crazy business impact

— Vladimir Keil, a16z Podcast
Key Insight
Boring is itself a moat: a function nobody wanted to own has drawn hundreds of tools that only made the old workflow faster and never changed how the work is done, so no incumbent ever reinvented it. The emotional pain is what gets a buyer in the room, and the P&L leverage (1% saved rivals 10% more sales) is what gets the CFO to sign.