Case Study: Intel achieves 123x BERT inference performance improvement with Numenta

A Numenta Case Study

Preview of the Intel Case Study

Intel achieves 123x BERT inference speedup with Numenta

The customer, Intel, faced the challenge of meeting the high-throughput and low-latency performance demands required for real-time natural language processing applications like Conversational AI. The size and complexity of Transformer models made this nearly impossible to achieve cost-effectively.

Numenta implemented a solution by combining its proprietary, neuroscience-based technology with Intel's new 4th Gen Xeon Scalable processors featuring Advanced Matrix Extensions. This integration into Intel's OpenVINO toolkit resulted in a 123x throughput performance improvement for BERT-Large Transformers while achieving sub-9ms latencies. Numenta's solution provided a highly scalable and cost-effective option for running large deep learning models, opening new possibilities for deploying Transformer models in production for time-sensitive AI applications.


View this case study…

Numenta

4 Case Studies