Toloka
8 Case Studies
A Toloka Case Study
Shopify needed to ensure its Sidekick AI commerce agent maintained high accuracy as it rapidly deployed new specialized skills for merchants. The challenge was generating a reliable quality signal to catch subtle failures in agent output that could erode merchant trust, and to do so at the scale and pace of new skill releases. To address this, Shopify partnered with vendor Toloka to build a ground-truth generation system.
Toloka implemented a human-in-the-loop solution where its experts annotated failing conversations to create golden datasets for fine-tuning. This process improved a key skill-judge's accuracy from 75% to 89% in a month through prompt engineering alone. The solution provided direct model improvement through supervised fine-tuning, delivered actionable analytics on root failure causes, and cut per-task annotation costs in half through efficiency gains.