The Art of Choosing the Right Model. Why We Created the Commerce & Enterprise Benchmark (CEB) 

Choosing an AI model for production systems rarely is just a matter of down to a simple comparison of charts on public leaderboards. Vendor marketing promises and academic tests rarely answer the key questions: how a given LLM will handle our specific tasks, how much it will cost to execute them, and what the quality of the output will be. Evaluation only becomes valuable when it reflects the actual environment in which the system is supposed to operate, tackling problems and datasets aligned with our needs. 

The Commerce & Enterprise Benchmark (CEB for short) addresses these challenges in a single evaluation suite. It provides a reproducible way to test the capabilities and limitations of AI systems using tasks relevant to Commerce & Enterprise solutions. 

Why did we create our own AI benchmark? 

CEB is a tool we created for our own needs to reproducibly test the real-world capabilities of models, agents, and evaluation harnesses in our daily engineering work at Univio. Today, we are opening it up and sharing our insights publicly. 

A large portion of existing benchmarks relies on synthetic data and public datasets. Unsurprisingly, models are trained on this exact same data. In many cases, models rank high on leaderboards simply because they memorized the questions during training (i.e., data contamination, meaning the leakage of test data into training sets). As a result, public benchmarks often measure a model’s memory rather than its reasoning capabilities. 

A stellar score in SWE-bench, MMLU, or GSM8K does not translate to whether the model will correctly decline a Polish surname, work efficiently with Jira and HubSpot, or maintain document consistency after tens of edits. 

Public leaderboards and CEB operate in entirely different domains. The former frequently measure universal academic competencies. We measure the models’  effectiveness in solving problems encountered in e-commerce and enterprise organisations.

What is CEB? 

It is a closed set of tasks selected from real-world problems encountered by Univio employees and our partners. The test architecture is based on Inspect AI from the UK AI Safety Institute. 

Coding tasks are run in isolated sandboxes. 

What exactly do we evaluate? 

We apply a range of evaluation dimensions: 

  • Linguistic tasks: declension, spelling, punctuation, and vocabulary comprehension, verified independently for Polish, English, and German. 
  • Long context and editing consistency: multimodal operations and handling tasks that exceed the allowable context window size. 
  • Working with data, business tools, and technical integrations: e.g., Jira, Confluence, HubSpot, Excel. 
  • Working with structured data and data contracts. 
  • Fact validation: knowledge tests identifying the model’s knowledge cutoff and specialized domain expertise. 
  • Coding and automation: currently analyzed in JS, Python, and PHP. 

Thanks to the private nature of the benchmark, our test data does not end up in public training corpora. We do not use synthetic data; we rely on real problems that have been anonymized and encapsulated into test cases. 

In addition to scoring the results, we also calculate the number of tokens used and their cost. We record the execution time of the test, but we do not factor it into the rankings—we don’t consider it a reliable metric when inference is handled over a network.  

Our goal is an objective measurement of a model’s true competencies in commercial applications. 

What does CEB not do? 

We do not treat our benchmark as a universal ranking for every use case. We are also aware of its current limitations; for example, the task database for the German language is currently too small to draw conclusive results, but we are constantly expanding it. 

We are also not looking for a perfect model for everything. We are looking for optimal solutions for specific use cases and budgets. 

 We regularly test AI model performance across various use cases and sectors. Curious how the ones you use at work are doing? We’d gladly compare our findings with your industry and needs. Contact us! 

Our Experts
/ Knowledge Shared

Expert Knowledge
For Your Business

As you can see, we've gained a lot of knowledge over the years - and we love to share! Let's talk about how we can help you.

Contact us

<dialogue.opened>