Cerebras
17 Case Studies
A Cerebras Case Study
Tavus, an AI video research company, faced a challenge in creating a realistic conversational experience for its digital twin video platform. The latency from its large language model (LLM) was the slowest part of its pipeline, causing delays that made interactions feel unnatural. Tavus needed to minimize both the Time to First Token (TTFT) and improve the Token Output Speed (TPS) to achieve a seamless, real-time feel for its Conversational Video Interface.
By integrating Cerebras's fast inference engine for the Llama 3.1-8B model, Tavus achieved a 66% reduction in TTFT and a 300% increase in TPS. This solution from Cerebras resulted in an overall LLM latency reduction of 100-1000%, dramatically improving the responsiveness of Tavus's system and enabling a more lifelike and immediate user experience.