AI Factory Insider

Pradeep Gupta & Stephen Jones on turning CUDA libraries into agent skills

Pradeep Gupta & Stephen Jones· VP of Industry Solutions and CUDA Architect at NVIDIA at NVIDIA
·~39 min·English·NVIDIA
AI InfrastructureGPUAgentsOpen Source
TL;DR

Two NVIDIA veterans explain how CUDA stays vertically integrated yet horizontally open, and why turning its 900+ libraries into agent skills reshapes the developer's job.

01Core Framing

Vertically Integrated, Horizontally Open

NVIDIA tunes the whole stack top-to-bottom for performance, then opens every layer so partners can plug in their own hardware, software, data, and models.

We want to open each and every layer to our ecosystem so that everyone can go and be the part of this entire platform, which is what we call an accelerated computing platform.

— Pradeep Gupta, AI Factory Insider
Key Insight
The phrase resolves an apparent contradiction: a tightly co-engineered stack usually implies lock-in, but NVIDIA's pitch is that vertical tuning and horizontal openness are separable choices made layer by layer, which is how one platform claims to serve seventeen verticals and, in Pradeep's framing, a hundred-trillion-dollar span of industries.

02The Foundation

What CUDA Actually Is

CUDA is the layer that carries any application down to the GPU, and its secret is that it is not one thing but many stacked layers, from frameworks at the top to the driver at the bottom.

CUDA is the magic that connects your application to the GPU hardware. Anything that you are doing that is GPU accelerated, that involves accelerated computing, it goes through CUDA.

— Stephen Jones, AI Factory Insider
Key Insight
Calling CUDA the magic understates a deliberate design: by making the stack deep and modular rather than monolithic, NVIDIA lets a developer touch it at any level and still reach the hardware, which is also what makes the platform hard to replace.

03The Library Layer

The Car and Its Wheels

On top of CUDA sit more than nine hundred CUDA-X libraries, reusable pieces that depend on one another so you extend proven work instead of reinventing it.

I kind of think of it in terms of CUDA as like a whole car it's a whole machine and the libraries are the wheels right nobody can use the car without the wheels. But you wouldn't reinvent the wheel either.

— Stephen Jones, AI Factory Insider
Key Insight
The interoperability is the real moat: cuDNN leans on cuBLAS, which leans on CUTLASS, so an AI workload inherits twenty years of scientific-computing optimization for free, and every new library deepens the dependency web that keeps work on the platform.

04Demystifying Open

Open Does Not Mean Open Source

NVIDIA's open means CUDA is freely available to take, use, and build a business on without permission, not that it is open source or vendor-neutral, though many libraries and even Nemotron are open source too.

Open means it's freely available. You can go and build your businesses on top of that.

— Pradeep Gupta, AI Factory Insider
Key Insight
The redefinition does strategic work: NVIDIA separates permission to use from source availability and hardware portability. It does commit to open standards and contributes to open-source projects, yet only the freely-available promise is unconditional while CUDA itself stays tied to NVIDIA GPUs, which is how 'open' wins developer goodwill without loosening the anchor.

05Hardware-Software Co-Design

Half His Time on the Other Side

A CUDA software architect spends fully half his time with the hardware team, so each generation from Turing to Vera Rubin is co-designed, and new CUDA advances are pushed back to older GPUs like Ampere and Turing.

I'm in the software organization as one of the architects of CUDA, but fully half my time is spent working with the hardware team.

— Stephen Jones, AI Factory Insider
Key Insight
Co-design makes backward-compatible software gains part of the platform: because one team shapes both sides, where older hardware can run them new CUDA advances extend a GPU's useful life after purchase, which strengthens the incentive to stay on NVIDIA rather than switch.

06The New Interface

Every Library Becomes an Agent Skill

NVIDIA is exposing every CUDA-X library as an agent skill, so an agent can pull in the right library for a job, which Pradeep expects to lift adoption by an order of magnitude.

think of CUDA skills to every CUDA library is actually going to drive the adoption of this to, I will say, an order of magnitude higher what it has been done today.

— Pradeep Gupta, AI Factory Insider
Key Insight
Skills are documentation written for agents instead of humans: they carry not just what a library does but when and why to choose it. If millions to billions of agents each call many skills, the library easiest for an agent to discover wins, so packaging, not only performance, becomes a distribution strategy.

07The Developer's New Job

You Manage the Agents

Developers are shifting from writing code to writing prompts and guiding the agents that write it, with Stephen reporting he can now do about three times more.

I must admit that I haven't written code, mostly this year, because I now write prompts, because the agents write the code.

— Stephen Jones, AI Factory Insider
Key Insight
The role inverts from author to manager: Stephen already leads agents as a team and Pradeep calls everyone an agents manager. The scarce skills become deep domain knowledge worth codifying into a skill and full-stack system thinking about how to run fleets of agents efficiently, not hand-writing the kernels.