Case Study: EliseAI achieves lower latency, lower inference costs, and frontier-matching accuracy with Baseten

A Baseten Case Study

Preview of the EliseAI Case Study

EliseAI cuts inference latency 80% with Baseten

EliseAI, an AI startup specializing in housing and healthcare operations, faced significant challenges with closed-source AI models. Their real-time conversational agents, which handle tasks like scheduling and maintenance requests, required sub-second latency and high accuracy. The vendor's closed-source APIs resulted in high costs, slow response times of around 2.2 seconds, and a lack of controllability, preventing the launch of a critical voice agent feature.

By partnering with Baseten, EliseAI used Baseten Training and research support to fine-tune smaller, open-source models like Qwen-4B. Baseten's team employed advanced techniques such as On-Policy Self-Distillation to boost accuracy. The solution led to a 60% reduction in inference cost, a reduction in p90 latency from 2.2 seconds to 250ms, and 99% accuracy on key tasks. This allowed EliseAI to successfully deploy their latency-sensitive voice agent and own their AI infrastructure.


View this case study…

Baseten

23 Case Studies