Lightcone

Jeffrey Morgan on Why Open Models Win the Tokens, Not the Budget

Jeffrey Morgan· Co-founder and CEO of Ollama at Ollama
·~57 min·English·Y Combinator
Open SourceInferenceAgentsAI Company
TL;DR

Ollama's CEO explains why enterprises are moving the majority of their tokens to open models — cost gets them in the door, control keeps them — and why the durable value is the glue layer that turns a fragmented universe of models, harnesses and hardware into something that just works.

01What He's Seeing

Open Models Took Over the Enterprise

Enterprises are shifting the bulk of their AI workload to open models, and coding agents are what tipped it.

I think the biggest thing we're seeing is a shift to open models, especially in enterprise

Jeffrey Morgan, Lightcone
Key Insight
The trigger wasn't a price cut — it was capability. Open models only crossed into enterprise once they were good enough to run coding and agentic workloads end to end, which is also the most token-hungry work there is.

02Why They Adopt

Cost Is the Door, Control Is the North Star

Cost is what pulls enterprises to open models, but owning and customizing the model is the reason they stay.

Cost is by far the largest pain point that open models can jump in and solve but you know every business has a vision of getting better control over AI and customizing it for their business and that's really their north star.

Jeffrey Morgan, Lightcone
Key Insight
Framing cost as the entry and control as the goal reframes the whole buying decision: a lab that only competes on price is fighting over the doormat, while the durable relationship is with whoever lets the business shape the model to its own data and needs.

03The Economics

Most of the Tokens, a Sliver of the Bill

Open models will carry the large majority of an enterprise's tokens while taking only a small share of its spend.

The super majority of tokens and this is our take it will be open models within a business.

Jeffrey Morgan, Lightcone
Key Insight
These two numbers point in opposite directions on purpose. Because open tokens are so cheap, a business can route 80–90% of its volume through them yet still send most of its dollars to frontier labs for the hard 10% — so token share and revenue share stop being the same story.

04How They Coexist

A Router, Not a Winner

The steady state isn't open beating closed — it's a router sending hard work to frontier models and everything else to open ones.

a lot of the kind of line item work can happen through open models and the collaboration of the two together.

Jeffrey Morgan, Lightcone
Key Insight
The interesting tension is who owns the router. Whoever decides which model handles a request controls the spend and the customer relationship — which is exactly why the frontier labs would rather you not have one.

05What Ollama Is

The Operating System for Open Models

Ollama's real job is to be the glue that makes any harness run any model on any hardware.

I do think the OS, which is generally a cliche analogy to use, is a good one because you've got the drivers for the hardware and the providers and the inference layer, but you also have the application runtime and making sure that the harness works.

Jeffrey Morgan, Lightcone
Key Insight
The OS analogy is also a moat claim. Operating systems win by owning the messy integration nobody else wants to maintain — and Morgan is betting the same combinatorial glue work between models, harnesses and chips is where a durable business hides in open source.

06Operating at Scale

Day Zero Is a Fire Drill

Making a new open model usable on launch day means packaging harness, model and hardware together — usually in the last 24 hours.

A lot of this stuff comes together in the last 24 hours before the model gets released and so it's generally a fire drill.

Jeffrey Morgan, Lightcone
Key Insight
A model's weights are the easy part. The work that decides whether a launch lands — harness support, tool-calling quirks, enough capacity, hardware tuned for speed — is invisible integration, and doing it reliably on a 24-hour clock is itself the product.

07What's Coming

The Return of Unlimited Tokens

A new class of ultra-cheap flash models, chained together, is bringing back the feeling of not having to count tokens.

this new class of flash models where they're good enough for 80% of the tasks, they're really fast and they're ultra cheap.

Jeffrey Morgan, Lightcone
Key Insight
This quietly settles the old “one giant god model” debate. When cheap models are good enough for 80% of tasks and can be chained, the winning pattern becomes orchestration over raw scale — and cost stops being the thing a developer has to ration.

08The Origin

Lost in the Wilderness

Ollama came after two years of pivots that went nowhere; the unlock was shipping the first version in two weeks instead of overthinking it.

And before that was two years of just frankly overthinking the customer, the product, and just not getting something out there.

Jeffrey Morgan, Lightcone
Key Insight
The lesson isn't that the wilderness years were wasted — the team learned what good looked like at Docker and VMware. It's that knowing the shape of the problem meant nothing until they forced themselves to ship, and the two-week version beat two years of planning.