AI Briefing
KO

How NVIDIA's Inference Software Stack Is Lowering Token Costs

·2026.07.01 00:00

Key point

NVIDIA is dramatically lowering the cost per token in AI inference through its Blackwell platform and optimized software stack.

1 / 2

Details

As enterprises move beyond AI pilot stages into actual production, the criteria for infrastructure decisions are shifting from simple chip specifications to Cost per Token. This means how many useful tokens can be delivered per dollar, power, and Latency.

NVIDIA's full-stack inference software is co-designed with GPUs, CPUs, networking, and systems to continuously improve hardware performance. In particular, on the NVIDIA Blackwell platform, the software stack reduced the token cost of the DeepSeek V4 model by up to 5x in just one month.

Adoption cases from major companies are as follows:

  • Baseten: Improved DeepSeek V4 Pro's tokens per second generation by up to 50% using TensorRT-LLM.
  • Cognition: Scaled reinforcement learning workloads without building infrastructure through the NVIDIA Dynamo framework.
  • Deep Infra: Instantly serving the latest open source models including DeepSeek V4 based on Blackwell.
  • Together AI: Accelerated model optimization for Cursor's real-time coding experience using TensorRT-LLM.

Existing web/SaaS workloads had predictable patterns, but Agentic AI is different. Agents perform reasoning, planning, tool calling, sub-agent creation, and more, requiring stateful workflows distributed across data centers to be executed, which calls for a new economic approach.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.