Hashlogics
Glossary

What is an LLM benchmark?

A model can top the leaderboard and still misread every invoice in your queue. The exam and your documents are not the same test.

LLM benchmark

model benchmarks

An LLM benchmark is a fixed set of questions or tasks used to score and compare language models, producing a single number or leaderboard rank. It measures performance on that test set, not on any particular business task.

MMLU is the one buyers hear about most. It stands for Measuring Massive Multitask Language Understanding, and it asks a model multiple-choice questions across dozens of subjects, from law to elementary maths. A score is the percentage answered correctly.

Other benchmarks target narrower skills. HumanEval scores generated code against unit tests. GSM8K scores grade-school maths word problems. Each one is a proxy for a capability. It gets picked because it is cheap to grade at scale, not because it matches your task.

Why it matters

Rank on a leaderboard is not accuracy on your task

A benchmark question is written once, graded automatically, and answered by every model that gets tested. Your task is none of those things. It runs on your documents, your formatting and your edge cases. The grading rule is whatever your business actually needs to be true.

Contamination makes the gap worse. Benchmark questions are public, so they end up in the training data of later models. A model can then score well by having memorised the answer, not by reasoning its way there. A published score can be real and still not travel.

Two models close in overall rank can differ sharply on one narrow skill, and that narrow skill can be the exact one your product needs. Extracting a dollar figure from a scanned invoice is nothing like answering a bar exam question, even though both get called reasoning.

  • 01Treat a benchmark score as a shortlist filter, not a final answer.
  • 02Watch for contamination: a suspiciously high score on an old, public benchmark is a reason to test harder, not trust more.
  • 03A model that wins on MMLU can still lose on your specific document type.
From leaderboard to your taskLive
  1. LeaderboardPublic rank, fixed questions.
  2. ShortlistTwo or three candidate models.
  3. Task evalYour documents, your grading rule.
  4. Score gapRank and task accuracy can diverge.
  5. DecisionPick the model that wins your eval.

The step teams skip is the third one. Benchmark rank picks a shortlist; only a task-specific eval picks the model.

Questions, answered

Common questions

01What does MMLU measure?

MMLU tests a model on multiple-choice questions spanning around 57 subjects, from history to professional law to abstract algebra. The score is the percentage answered correctly, and it is meant as a broad check of general knowledge and reasoning, not a specific skill.

02Do benchmarks predict real performance?

Only loosely. A high benchmark score shows a model can handle a wide range of general questions. It says little about accuracy on your specific documents, format or edge cases. Public benchmarks can also be contaminated by appearing in later training data, which inflates scores without a matching gain in ability.

03How is a benchmark different from an eval?

A benchmark is public, fixed and shared across every model that gets tested. An eval is built by a team for one task, using their own real data and their own definition of correct. A benchmark ranks models in general; an eval ranks them for you.

04Should I pick a model based on benchmark rank alone?

Use rank to build a shortlist of two or three candidates. Then test each one against a small set of your own graded examples before deciding. The model that wins the leaderboard does not always win the task.

Written by Abdul Basit, CEO, HashlogicsVerified
Start

Let’s build the one that runs after.

We build AI agents and automation, then stay on under an agreed service level. A senior engineer reads every brief, and your call gets scheduled within 24 hours.

What happens next

  1. 01

    You send a brief or book a call

    Two minutes, whichever you prefer.

  2. 02

    A senior engineer replies within 24 hours

    Not a sales rep.

  3. 03

    Honest scoping, in writing

    And if we’re not the right fit, we say so.

Abdul Basit, CEO of Hashlogics

“I started Hashlogics because too many teams ship a demo, get paid, and disappear. We build to a standard we’d run ourselves — and we stay to keep it running.”

Abdul Basit · CEO · a direct line

Not ready to talk? Take the checklist.

12 questions to ask any AI agency before you sign. They separate a demo shop from a team that ships to production.

Get the checklist

Free · no newsletter