Case Study: Together AI boosts inference speed and reduces latency with WEKA NeuralMesh

A Weka Case Study

Preview of the Together AI Case Study

Together AI reduces inference latency for 500,000+ developers with Weka

Together AI, a company advancing open-source artificial intelligence, needed to reduce latency and deliver ultra-fast data access for its industry-leading inference engine across its massive GPU-powered cloud environment. To support its rapid growth and serve its community of over 500,000 developers, it turned to WEKA for a solution, leveraging WEKA's NeuralMesh.

WEKA implemented its Augmented Memory Grid capability to reduce the time involved in prompt caching and improve the flexibility of leveraging this cache across multiple nodes. This solution successfully reduced latency for Together AI, benefiting its vast developer community and contributing to the company's ability to provide the fastest inference speeds in the industry.


View this case study…

Weka

44 Case Studies