Hashlogics
Glossary

What are AI evals?

Without them you edit one line of a prompt, and six weeks later you learn the system has been reading the wrong column on every payroll file.

AI evals

EvaluationsLLM evals

AI evals are a graded set of test cases run against a model, scoring its output against a known correct answer. They tell a team whether a prompt, model or data change made the system better or worse.

An eval has three parts. A dataset of inputs with the answer you expect. A grader that decides whether the output matches. A score you can compare against last week's run.

Graders vary by task. Pulling a gross payroll figure out of a register has one right answer, so an exact match works. Summarising a claim file does not. There you score against a rubric, or hand it to a second model told what good looks like, and sample the results yourself.

Why it matters

A prompt is a dependency you cannot see

Ordinary software fails loudly. A model fails quietly and stays plausible. Change one instruction to fix a rare document type, and the fix can quietly cost you accuracy on the common one. Nothing throws. The output still looks right.

Evals turn that invisible change into a number that moved. They also answer a question a regulated buyer will ask: how do you know it still works? Pointing at a passing suite with dates on it is a better answer than describing your process.

This matters most in document-heavy insurance work. Reconciling a payroll register against an IRS Form 941 has a checkable right answer, which means accuracy is measurable and drift is catchable.

  • 01Run them on every prompt change, model version bump and retrieval tweak.
  • 02Keep the failures. A case the system got wrong once is the most valuable test you own.
  • 03Score the cases you would be embarrassed to get wrong, not the easy majority.
The eval loopLive
  1. CollectReal inputs, not invented ones.
  2. LabelA human sets the right answer.
  3. RunModel output, every change.
  4. GradeExact match, or a rubric.
  5. GateScore drops, release stops.
  6. AddEvery new failure joins the set.

The gate is the part teams skip. A suite that reports a score but never blocks a release is a dashboard, not a test.

Questions, answered

Common questions

01How are evals different from unit tests?

A unit test asserts one exact output and fails if a single character differs. An eval scores a distribution, because the same input can produce different wording each run. So evals report a rate across many cases rather than a pass or fail on one, and the gate is a threshold you set.

02How many test cases do you need?

Fewer than teams expect, if they are the right ones. A few dozen genuinely difficult, correctly labelled cases catch more regressions than a thousand easy ones. Labelling cost is the real constraint, since a human has to decide the right answer for each.

03Can a model grade its own output?

For open-ended tasks, yes, with checks. A separate model given a clear rubric grades summaries and answers reasonably well, and it is far cheaper than a human. Sample its verdicts by hand regularly, because a grader can drift the same way the system under test can.

04Do evals prove the system is compliant?

No. Evals prove accuracy and catch drift. Compliance also needs logged inputs and outputs, versioned prompts, recorded human overrides and a defined escalation point. Evals are one of the artefacts an examiner finds convincing, not the whole file.

Verified
Start

Anyone can ship the agent. We answer the pager.

We build AI agents and automation, then stay on under an agreed service level. A senior engineer reads every brief, and your call gets scheduled within 24 hours.

What happens next

  1. 01

    You send a brief or book a call

    Two minutes, whichever you prefer.

  2. 02

    A senior engineer replies within 24 hours

    Not a sales rep.

  3. 03

    Honest scoping, in writing

    And if we’re not the right fit, we say so.

Abdul Basit, CEO of Hashlogics

“I started Hashlogics because too many teams ship a demo, get paid, and disappear. We build to a standard we’d run ourselves — and we stay to keep it running.”

Abdul Basit · CEO · a direct line

Not ready to talk? Take the checklist.

12 questions to ask any AI agency before you sign. They separate a demo shop from a team that ships to production.

Get the checklist

Free · no newsletter