Case Study: Baseten achieves faster, lower-cost AI inference with NVIDIA Blackwell, Dynamo, and TensorRT-LLM

A NVIDIA Case Study

Preview of the Baseten Case Study

Baseten boosts AI inference throughput 5x with NVIDIA Blackwell

The customer, Baseten, faced the challenge of scaling its AI inference platform to meet growing demand for large, complex models, particularly reasoning models that require significant compute and memory. To help its customers scale quickly, Baseten adopted NVIDIA's latest data center GPU architecture, NVIDIA Blackwell, on Google Cloud, along with the NVIDIA Dynamo inference framework and NVIDIA TensorRT-LLM.

The solution implemented with NVIDIA delivered breakthrough performance and efficiency. Baseten achieved a 5x higher throughput for high-traffic endpoints, a 2x better price-performance when serving frontier reasoning models, and up to a 38% reduction in latency for serving the largest LLMs. This resulted in dramatically lower costs and a significantly improved user experience for their customers.


View this case study…

NVIDIA

18 Case Studies