Case Study: Shopify boosts skill-agent accuracy at scale with Toloka

A Toloka Case Study

Preview of the Shopify Case Study

Shopify boosts skill-judge accuracy from 75% to 89% with Toloka

Shopify needed to ensure its Sidekick AI commerce agent maintained high accuracy as it rapidly deployed new specialized skills for merchants. The challenge was generating a reliable quality signal to catch subtle failures in agent output that could erode merchant trust, and to do so at the scale and pace of new skill releases. To address this, Shopify partnered with vendor Toloka to build a ground-truth generation system.

Toloka implemented a human-in-the-loop solution where its experts annotated failing conversations to create golden datasets for fine-tuning. This process improved a key skill-judge's accuracy from 75% to 89% in a month through prompt engineering alone. The solution provided direct model improvement through supervised fine-tuning, delivered actionable analytics on root failure causes, and cut per-task annotation costs in half through efficiency gains.


View this case study…

Toloka

8 Case Studies