Baseten
23 Case Studies
A Baseten Case Study
Hebbia, a provider of AI-powered intelligence for leading financial institutions, faced a challenge in delivering the fast, reliable, and low-latency AI inference required by its finance professional user base. Their previous provider could not offer the flexibility or performance needed for bursty, latency-sensitive chat traffic, which was critical for their customers' high-stakes decision-making.
By partnering with Baseten and utilizing its inference stack for a dedicated open-source LLM deployment, Hebbia gained a high-performance solution. Baseten implemented techniques like KV-aware cache routing and speculative decoding to significantly reduce latency. The results included a 2.5x improvement in tokens per second, a 4x improvement in time to first token, and a more than 10x reduction in cost, all while meeting the demanding reliability standards of their enterprise customers.