Case Study: Databricks builds the OfficeQA benchmark with SuperAnnotate

A SuperAnnotate Case Study

Preview of the Databricks Case Study

Databricks builds OfficeQA with SuperAnnotate for 90,000-page benchmark and 246 questions

Databricks faced the challenge that standard AI benchmarks were inadequate for evaluating how well large language models perform on real-world business tasks with messy, complex corporate data. To address this, they partnered with SuperAnnotate to create a more realistic benchmark that tests grounded reasoning, the ability to find and understand information within difficult documents.

SuperAnnotate provided a managed workforce of subject matter experts through its SME Careers service to build the OfficeQA benchmark. This involved creating and validating challenging questions based on 90,000 pages of U.S. Treasury Bulletins. The results revealed a significant performance gap: even top AI models achieved less than 45% accuracy on the benchmark, proving that standard tests are insufficient and highlighting the critical role of human expertise in developing trustworthy business AI.


View this case study…

SuperAnnotate

27 Case Studies