Hashlogics
Glossary

What is LLM-as-judge?

Without an audit step, the judge just agrees with whichever answer sounds most confident, and nobody notices until a customer does.

LLM-as-judge

Model-graded evaluation

LLM-as-judge is a method for scoring a language model's output. A second model reads a rubric and the answer, then grades the result. It replaces manual review for tasks too open-ended for exact-match scoring, such as a summary or a support reply.

A judge prompt names what good looks like. Did the reply answer the question? Does it invent a fact the source document never stated? Is the tone inside the allowed range? The judge model reads the output against those criteria and returns a score, sometimes with a reason attached.

This only works where exact match does not. A payroll figure has one right answer, so you compare strings. A support reply or a research summary has many acceptable phrasings instead, and that is the gap LLM-as-judge fills.

Why it matters

It scales evaluation, and it scales its own mistakes

A human can grade a few hundred outputs a week, carefully. A judge model can grade thousands overnight. That difference is why teams reach for it as soon as an eval set grows past what one reviewer can carry.

The catch is that a judge inherits the biases of any language model. Researchers have a name for one: self-preference. A judge rates output from its own model family higher than equally good output from a different one. Length bias shows up too. A longer answer often scores better, regardless of whether the extra words add anything.

None of that disqualifies the method. It means the judge is a component you test, not a source of ground truth you trust on install. The discipline is the same one that makes the system under test defensible. Measure it against something you already know is right.

  • 01Audit the judge against human-labelled examples before its scores count for anything.
  • 02Track judge-human agreement as its own number, and re-check it after every model swap. A rubric with named criteria beats an open "rate this from 1 to 10".
Standing up a judge you can trustLive
  1. RubricNamed criteria, not a vague score.
  2. Human labelsA sample graded by a person first.
  3. Judge runSame sample, scored by the model.
  4. Agreement rateHow often judge and human matched.
  5. GateBelow threshold, the judge is not trusted yet.
  6. RecheckEvery prompt or model change, again.

Skip the agreement check and you have not built an eval system. You have automated a guess.

Questions, answered

Common questions

01Is LLM-as-judge reliable?

It is reliable only after you check its agreement with human graders on a sample. Confirm the rate is high enough for the decision it supports. An unaudited judge can be confidently wrong, since it scores fluent, well-formatted answers higher regardless of accuracy. Treat the audit as mandatory, not optional.

02Should the judge be a different model from the one being graded?

Using a different model family reduces self-preference bias, where a judge rates its own family's phrasing higher. It does not remove the need to audit the judge separately. A stronger, differently-sourced model is a reasonable default, still checked against human labels before its scores are trusted.

03What tasks should never use LLM-as-judge?

Anything with one checkable correct answer, such as a dollar figure or a database lookup, should use exact-match scoring instead. A judge model adds cost and a new failure mode without adding accuracy there. Reserve it for open-ended output where a rubric is the only realistic grading method.

Written by Abdul Basit, CEO, HashlogicsVerified
Start

Let’s build the one that runs after.

We build AI agents and automation, then stay on under an agreed service level. A senior engineer reads every brief, and your call gets scheduled within 24 hours.

What happens next

  1. 01

    You send a brief or book a call

    Two minutes, whichever you prefer.

  2. 02

    A senior engineer replies within 24 hours

    Not a sales rep.

  3. 03

    Honest scoping, in writing

    And if we’re not the right fit, we say so.

Abdul Basit, CEO of Hashlogics

“I started Hashlogics because too many teams ship a demo, get paid, and disappear. We build to a standard we’d run ourselves — and we stay to keep it running.”

Abdul Basit · CEO · a direct line

Not ready to talk? Take the checklist.

12 questions to ask any AI agency before you sign. They separate a demo shop from a team that ships to production.

Get the checklist

Free · no newsletter