Case Study: Scaled Cognition achieves 114% higher LLM accuracy with Weights & Biases

A Weights & Biases Case Study

Preview of the Scaled Cognition Case Study

Scaled Cognition builds reliable LLMs with 114% higher accuracy using Weights & Biases

Scaled Cognition builds highly reliable AI, including the APT-1 LLM, for customer support in regulated industries like banking and healthcare. Their challenge was that general-purpose LLMs were unsuitable as they often fabricated responses and failed to enforce company policies, a critical issue in fields requiring absolute accuracy. To address this, they needed a robust training and monitoring solution, which led them to use the Weights & Biases platform.

Using the Weights & Biases Python SDK, Scaled Cognition integrated a low-friction observability layer into their training pipeline. This allowed them to track custom metrics, monitor system resources like GPU usage in real-time, and maintain a reproducible record of all experiments. The solution from Weights & Biases empowered them to build APT-1, which achieved a 114% higher accuracy at Pass⁵⁰ compared to the best general-purpose LLM and delivers 100% reliability on critical tasks with no performance decay.


View this case study…

Weights & Biases

49 Case Studies