Labelbox
51 Case Studies
A Labelbox Case Study
Meta Superintelligence Labs needed a benchmark that could evaluate advanced AI reasoning as traditional LLM tests became saturated. They required tasks grounded in practical reasoning with detailed rubrics for partial credit and robust quality control. To address this, they partnered with the vendor Labelbox for data foundation services.
Labelbox produced the expert-authored data foundation for the Grounded Integration Measure (GIM) benchmark, creating 820 problems with structured scoring and quality assurance. This enabled Meta to calibrate a model over 200,000 prompt-response pairs. The result was the release of a durable benchmark, GIM-615, which showed 20% of items were above frontier model ability and was used to evaluate 22 models across 47 configurations.