Case Study: Meta Superintelligence Labs builds a durable frontier AI reasoning benchmark with Labelbox

A Labelbox Case Study

Preview of the Meta Superintelligence Labs Case Study

Meta Superintelligence Labs builds GIM with Labelbox data and 820 expert-authored problems

Meta Superintelligence Labs needed a benchmark that could evaluate advanced AI reasoning as traditional LLM tests became saturated. They required tasks grounded in practical reasoning with detailed rubrics for partial credit and robust quality control. To address this, they partnered with the vendor Labelbox for data foundation services.

Labelbox produced the expert-authored data foundation for the Grounded Integration Measure (GIM) benchmark, creating 820 problems with structured scoring and quality assurance. This enabled Meta to calibrate a model over 200,000 prompt-response pairs. The result was the release of a durable benchmark, GIM-615, which showed 20% of items were above frontier model ability and was used to evaluate 22 models across 47 configurations.


View this case study…

Labelbox

51 Case Studies