Hashlogics
Answers

How to run evals on an AI agent before launch?

An agent you cannot score is an agent you cannot change. Evals are the difference.

Answered in short

5 things that decide this

  1. 01Build the eval set from real historical cases, including the messy ones, rather than examples written to make the agent look good.
  2. 02Agree the passing score before launch, while nobody is under pressure to ship, and treat that threshold as a release gate.
  3. 03Score each failure by type, because knowing an agent is 88% right tells you nothing about which 12% will embarrass you.
  4. 04Run the same suite on every prompt change and every model version, since that is what turns a model retirement into a routine afternoon.
  5. 05Keep a held-out set the builders never see, or the agent gets tuned to the test rather than the task.
Why teams skip this

The demo is not evidence

Most agents reach launch tested by hand. Somebody tries fifteen inputs, they look right, and everyone agrees it is working. That is a vibe, not a measurement.

The gap shows up later. A prompt gets edited to fix one complaint, something unrelated breaks, and nobody notices for a fortnight because there is no baseline to compare against.

This is also why model deprecation hurts teams without evals. Anthropic retired Claude Opus 4.1 on 5 August 2026, with the notice its policy commits to. Anyone holding an eval suite ran it against the replacement and moved on. Everyone else re-tested by hand and hoped.

  • A test suite is the thing that makes an AI system maintainable rather than frozen.
How we build one

Five steps that produce a usable suite

Collect real cases first. Pull actual inputs from your systems, weighted toward the awkward ones: incomplete records, unusual phrasing, the edge case that generated a complaint last year.

Agree the right answer with the person who owns the outcome. If two experienced staff disagree on a case, that disagreement is a finding, and the task may need narrowing before anything gets built.

Score by failure type, not one blended percentage. A wrong figure and a polite refusal to answer are different problems with different fixes.

Set the gate. Decide what score ships, what score blocks, and who may override it. Write it down while the decision is calm.

Automate the run so it happens on every change, because a suite that needs somebody to remember it stops being run by week three.

  • Hold back a portion of cases the builders never see. Tuning against a visible test set produces a score that flatters everyone and predicts nothing.
The eval gateLive
  1. Real casesPulled from your data, edge cases included.
  2. Agreed answersSigned off by the outcome owner.
  3. Typed failuresWrong, refused, and unsafe are not one number.
  4. ThresholdWritten down before launch day.
  5. Every changePrompts and model versions both.
  6. Held-out setNever shown to the people tuning.

This gate is why a model retirement is an afternoon rather than a crisis.

Questions, answered
01How many cases does a useful eval set need?

Enough to cover every failure mode you can name, which usually lands in the low hundreds rather than the thousands. Coverage of distinct situations matters more than volume. Fifty well-chosen awkward cases beat a thousand near-identical easy ones.

02Can another model grade the answers?

Model grading works for subjective qualities like tone and completeness, and it needs a human-checked sample to stay honest. For factual outputs, compare against the known correct answer instead. Never let the same model grade its own work without a human spot-check.

03When should the eval set be updated?

Add every production failure to the suite as it happens, so the same bug cannot ship twice. Review the whole set when the task changes or the data source shifts. A suite that never changes slowly stops describing the job.

Verified
Start

Anyone can ship the agent. We answer the pager.

We build AI agents and automation, then stay on under an agreed service level. A senior engineer reads every brief, and your call gets scheduled within 24 hours.

What happens next

  1. 01

    You send a brief or book a call

    Two minutes, whichever you prefer.

  2. 02

    A senior engineer replies within 24 hours

    Not a sales rep.

  3. 03

    Honest scoping, in writing

    And if we’re not the right fit, we say so.

Abdul Basit, CEO of Hashlogics

“I started Hashlogics because too many teams ship a demo, get paid, and disappear. We build to a standard we’d run ourselves — and we stay to keep it running.”

Abdul Basit · CEO · a direct line

Not ready to talk? Take the checklist.

12 questions to ask any AI agency before you sign. They separate a demo shop from a team that ships to production.

Get the checklist

Free · no newsletter