Make benchmarks true forms of data. Include revisions, provenance.

TPU and Wide EP

Fault tolerant EP in llmd

Back to costs. How much does it cost to generate 1m tokens in each of these distributed cases. What’s the cost for different pd settings.

Jensen keynote

Tldr

  1. AI factory based on tokens versus Data Centers based on files. Companies will be token producers and
  2. Openclaw framed as the gpt3 moment for agents. NemoClaw is Enterprise grade attached to a policy engine
  3. Frontier models for each industry. Presented as top of the Nemotron
  4. On robots, they led with robotaxis and a video that would have been novel 3 years ago

Olaf robot was awkward, timing off.

Showmanship on full display, total Super Bowl vibes.

raw notes

Bring customers to cloud

Broke out AI platform in context of cloud providers

Vertically integrated, horizontally open.

computing company. horizontal

Lead with Google Cloud

Accelerated Computing.
Application Acceleration

The Inference Inflection Arrives

Nvidia is the lowest cost, highest confidence for inference

1/3 of AI compute: anthropic, mls, multiple oss

Nvidia is drawing a clear line at the AI platform

Conflicting messages

  • vertically integrated, horizontally open

Inference is the ultimate hard

Dgxcloud to create kernels

2 charts: cost and performance

50x higher perf/watt
35x lower cost
atel - accused Jensen of sandbagging. He’s not wrong. Round of applause.

If you have the wrong architecture, even if it’s free, it’s not cheap enough.

Untouchable token cost due to increased codesign.

Fireworks, Together. 100x

Throughput and token speed for your AI factory

Tokens are a commodity.

Blackwell introduced a whole new tier

Spent time mapping tokens to revenues, teaching people how to monetize GPU

Highlighted the price perf
Grok benefits

Dynamo for PD to help Rubin and Grok work together. Prefill is the easy part.

Dynamo as the OS for AI factory

Called out challenges with Blackwell and NVL72. Going much smoother with Vera Rubin.

Blue field. Are we behind?

Every single SaaS company is going to become a gas company.

Wants to give employees a token budget.
Agents perceive reason and act