Case Study: Hebbia achieves faster, lower-cost AI inference with Baseten

A Baseten Case Study

Preview of the Hebbia Case Study

Hebbia improves TPS 2.5x with Baseten

Hebbia, a provider of AI-powered intelligence for leading financial institutions, faced a challenge in delivering the fast, reliable, and low-latency AI inference required by its finance professional user base. Their previous provider could not offer the flexibility or performance needed for bursty, latency-sensitive chat traffic, which was critical for their customers' high-stakes decision-making.

By partnering with Baseten and utilizing its inference stack for a dedicated open-source LLM deployment, Hebbia gained a high-performance solution. Baseten implemented techniques like KV-aware cache routing and speculative decoding to significantly reduce latency. The results included a 2.5x improvement in tokens per second, a 4x improvement in time to first token, and a more than 10x reduction in cost, all while meeting the demanding reliability standards of their enterprise customers.


View this case study…

Baseten

23 Case Studies