Case Study: Firmus Technologies achieves 6.5x higher long-context AI inference throughput with WEKA NeuralMesh and Augmented Memory Grid

A Weka Case Study

Preview of the Firmus Technologies Case Study

Firmus Technologies delivers 6.5x higher AI token throughput with Weka

Firmus Technologies, a sustainable AI infrastructure provider in Singapore, faced a challenge in delivering long-context AI inference at scale due to the hard memory ceiling of GPU HBM and host-local DRAM. This constraint limited context reuse, created operational tradeoffs, and resulted in lower performance and efficiency for their demanding workloads.

The solution was implemented using WEKA's Augmented Memory Grid on its NeuralMesh platform. This provided a shared memory tier that pooled capacity across hosts, breaking the per-node cache ceiling. The results included up to 6.5x higher sustained input token throughput, 34% lower time-to-first-token, and significantly improved power efficiency. For Firmus, this meant delivering more inference work without adding servers, aligning with their sustainability mandate to maximize every watt.


View this case study…

Weka

44 Case Studies