Case Study: Parallel Web Systems achieves 3x higher throughput and 2x lower latency with Baseten

A Baseten Case Study

Preview of the Parallel Web Systems Case Study

Parallel Web Systems boosts throughput 3x with Baseten

Parallel Web Systems, which builds infrastructure for AI agent web research, faced the challenge of scaling their multi-turn agentic pipelines. They needed high-throughput inference to handle tens of thousands of requests per hour and bursty traffic, but found the cost of using closed-source models to be unsustainable. They turned to vendor Baseten for a solution just before a major product launch.

Baseten first provided Parallel with its Model APIs to handle unpredictable launch traffic. As demand grew, Parallel moved to a dedicated Baseten deployment, where proprietary runtime optimizations like KV cache-aware routing and speculative decoding were implemented. The solution resulted in a 50% reduction in latency, a 3x improvement in throughput, and 3x cost savings compared to closed-source models. Baseten's infrastructure enabled Parallel to scale traffic 100x.


View this case study…

Baseten

23 Case Studies