Case Study: Speechify achieves real-time text-to-speech at scale with Baseten

A Baseten Case Study

Preview of the Speechify Case Study

Speechify cuts inference costs by 44% with Baseten

Speechify, a leading AI text-to-speech platform with over 60 million users, faced significant challenges managing its own inference infrastructure. Before partnering with Baseten, the complexity of their self-managed stack, which spanned 1,500 GPUs, slowed down model deployments and resulted in long cold starts. Their engineering priority was to reduce this operational overhead and improve latency for their new SIMBA 3.0 model to ensure the best user experience.

By moving its text-to-speech workloads to Baseten, Speechify implemented a solution using Truss for model deployments and leveraged Baseten's traffic-based autoscaling and multi-region inference. The results were substantial: Baseten helped Speechify achieve a 44% reduction in cost per million characters, a 30-50% decrease in p99 inference latency, and 4.5x faster cold starts. This allowed Speechify to retire 940 GPUs and free its engineering team to focus on product innovation.


View this case study…

Baseten

23 Case Studies