Prolific
56 Case Studies
A Prolific Case Study
Researchers at Honda Research Institute Europe, in collaboration with Bielefeld University, faced the challenge of creating a benchmark to test how well AI models understand vague temporal expressions like "recently" or "a long time ago." These phrases are subjective and difficult for AI to interpret without concrete human-defined standards. To build this benchmark, called TRAVELER, they needed high-quality, nuanced data on human perception, which led them to the participant platform Prolific.
The solution involved using Prolific to run detailed surveys that gathered human judgments on how vague time-related phrases should be interpreted. This human-validated data was integrated directly into their benchmark to create a probabilistic scoring system for testing AI. The results clearly showed a significant performance gap in AI models; accuracy dropped from 92% on explicit time references to just 45% on vague ones, demonstrating the measurable impact and success of the human-validated approach enabled by Prolific.