AI Briefing
KO

What Jensen Huang Said About Anthropic, OpenAI, China, and Inference Token Demand

·2026.04.17 09:00

Key point

Nvidia is solidifying its AI infrastructure dominance through supply chain, financial, and China strategies.

Details

Jensen Huang locates Nvidia's moat not in mere GPU performance, but in the relationships and trust that move the upstream supply chain. If the $250B in purchase commitments is the visible number, the real power lies in partners like SK Hynix, Micron, and TSMC believing in Nvidia's demand forecasts enough to expand capacity. Two years of pressure around CoWoS turned advanced packaging from a bottleneck into a standard, and Nvidia has effectively embedded its roadmap into the supply chain itself.

The only non-Nvidia training story Huang effectively acknowledges as an exception is Anthropic. He views Anthropic's use of TPU and Trainium not as a trend spreading to other labs, but as a special case that only came together by bundling capital together. The key point is that chip vendors are no longer simple suppliers, but are effectively acting as lenders of last resort, closing offtake deals alongside massive equity investments.

This logic extends to how Nvidia supports OpenAI, Anthropic, and neoclouds broadly. Huang put forward the principle of "do as much as needed, as little as possible," specifically naming CoreWeave, Nscale, and Nebius. Rather than taking on a cloud's P&L directly, Nvidia is choosing to design capital structures that generate demand while leaving the risk elsewhere.

The most emotionally charged moment was over China. When Dwarkesh Patel pressed that selling H20-class chips strengthens China's aggressive capabilities, Huang pushed back by emphasizing China's research talent, manufacturing capacity, and Huawei's growth. He believes that giving up the China market would mean surrendering the developer ecosystem and the expansion of CUDA, and ultimately the competition to set the standard for the global stack.

Finally, Huang suggested that the inference market is evolving not as a single uniform curve, but into a multi-layered market divided by latency and throughput. Nvidia is deliberately targeting the low-latency token segment where it can command higher ASPs, and differentiation is underway in which systems like Groq are being absorbed into CUDA. As such, modeling the inference token economy with a single price point or single curve misses the actual revenue structure.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.