Prolific
56 Case Studies
A Prolific Case Study
Cohere, an AI company, faced the challenge of reliably evaluating the performance of its large language models (LLMs) using human feedback, as it suspected that inherent human biases could skew the results and underrepresent critical factors like factuality. To investigate this, they partnered with the vendor Prolific to source participants for a detailed annotation study.
Using Prolific's platform, Cohere recruited qualified annotators and utilized a modified version of the open-source Potato annotation tool to conduct their experiments. The study revealed that human preference scores significantly underrepresent factuality errors and that annotators are less likely to identify these errors in outputs written in an assertive style. These findings, achieved through Prolific's services, provided Cohere with crucial insights into the biases present in human feedback, demonstrating that it is not a perfect gold standard for training or evaluating AI models.