Case Study: Krea achieves near-100% GPU utilization and ~40% MFU with WEKA NeuralMesh

A Weka Case Study

Preview of the Krea Case Study

Krea accelerates foundation model training with WEKA, reaching ~100% GPU utilization

Krea AI, an AI creative platform with over 30 million users, faced a critical challenge with its open-source Ceph storage system. The system could not handle the demands of training proprietary foundation models, suffering from severe metadata performance degradation, frequent failures, and inconsistent throughput. This resulted in extended GPU idle time, stalled research, and became a major constraint on the company's ability to innovate and release new models quickly.

By implementing the WEKA NeuralMesh data platform, Krea AI found a solution that provided POSIX compliance, reliable high-concurrency performance, and robust fault tolerance. The results were transformative, with WEKA enabling close to 100% GPU utilization on individual workloads and approximately 40% Model FLOPs Utilization on distributed runs. This eliminated training failures attributable to storage, accelerated model release timelines by weeks to months, and saved thousands of engineering hours, providing a scalable foundation for future growth.


View this case study…

Weka

44 Case Studies