**Notes on Jensen podcast with Dwarkesh

TLDR

I respect that Dwarkesh took a real run at 1) weakness of the CUDA moat, 2) threat of other chips like TPUs, and 3) Jensen’s desire to sell chips to China. Jensen wasn’t having any of it and took him to school on all points. Dwarkesh had an opportunity to press him more on each but always tacked away.

Feels like Dwarkesh and Dylan are working together and it’s clear that SemiAnalysis is pushing hard to expose weaknesses in the industry.

Most of this conversation was about the balance between enabling NVIDIA to continue to extend its ecosystem while not giving China the most advanced technology.

China is the largest contributor to OSS, open models. Built on the American tech stack. All 5 layers. US needs to go after all 5 layers. Point is every layer has to succeed. If we scare this country into AI is a nuclear bomb, I don’t know how you are helping the US.

Jensen: “Your premise is just wrong” x4 times

Interesting the Crusoe sponsored the podcast with a TTFT benchmark relative to vLLM

Jensen: Invest in all LFMBs – don’t go out of our way to pick winners. “It’s not our job to”

5 layer cake

  • Top Layer: AI Application
  • Bottom layer: Energy
    • US lags on energy

Jensen: AI models are developed are run best on NVIDIA

Raw notes

Pushed back hard on Dario’s nuclear metaphor. 

Instantaneous demand is better

Framed TPUs as a chip for tensors

NVIDIA has built CUDA 

Our market opportunity is a lot larger – because we support every application in the world.

Matrix Multiplies are important but not the only thing.

Hybrid SSM

You want an architecture that is generally programmable

60% of rev comes from big 5 hyperscalers

OpenAI has Triton - their own stack.

Replacements for Cuda. 

Cuda is a rich ecosystem – building on cuda first is incredibly smart.

Backend of triton - huge amounts

Triton

Verl, NeMO RL. Post Training. 

Flywheel is install base, programmatically of our architecture.

Dylan’s InferenceMax is sitting out there. TPU won’t come, Tranium won’t come. MLPerf - I would welcome trainum demonstrate their 40%. It makes no sense.

t33: Perf per dollar is so great – lowest cost token. Perf per watt is the highest in the world. Highest tokens per watt.**

I am the evidence. AI models are created on our stack run best on our stack.
You’ve got to be arguing something ridiculous first.

We are not a car. I am not a loser. Computing is not like a car. These ecosystems are hard to replace and most people don’t want to do it. Keep advancing the technology.

The industry is not a loser. That losing mindset doesn’t make sense to me.

You don’t have to move on - I’m enjoying it.

The crux - your arguments go to extremes. That if we give them any compute at all — we will lost everything. Those extremes are childish.

Is AI different?