Case Study: Tavus achieves real-time conversational video with Cerebras fast inference

A Cerebras Case Study

Preview of the Tavus Case Study

Tavus cuts LLM latency by 66% with Cerebras

Tavus, an AI video research company, faced a challenge in creating a realistic conversational experience for its digital twin video platform. The latency from its large language model (LLM) was the slowest part of its pipeline, causing delays that made interactions feel unnatural. Tavus needed to minimize both the Time to First Token (TTFT) and improve the Token Output Speed (TPS) to achieve a seamless, real-time feel for its Conversational Video Interface.

By integrating Cerebras's fast inference engine for the Llama 3.1-8B model, Tavus achieved a 66% reduction in TTFT and a 300% increase in TPS. This solution from Cerebras resulted in an overall LLM latency reduction of 100-1000%, dramatically improving the responsiveness of Tavus's system and enabling a more lifelike and immediate user experience.


View this case study…

Cerebras

17 Case Studies