Hashlogics
Glossary

What is eval-driven development?

Two engineers disagree about which prompt is better. Neither can prove it, so the loudest one wins and quality becomes a matter of taste.

Eval-driven development

Eval-first development

Eval-driven development is a way of building AI features where the scored test cases come first and the prompt comes second. Every later change is judged by whether the score went up or down, rather than by whether the output looks better.

The order is the whole idea. You collect real inputs, agree the right answer for each, and build the grader before anyone writes an instruction to the model. Only then does prompt work start.

Test-driven development is the obvious parallel, with one difference that matters. A unit test passes or fails. An eval reports a rate across many cases, so the gate is a threshold you choose and defend.

Why it matters

It replaces taste with a number

Prompt work without evals turns into an argument about wording. Someone tries a phrasing, reads five outputs, declares it better, and ships. Nobody notices that it got worse on the cases they did not read.

Building the test set first also forces a harder conversation early. Agree the right answer for fifty real inputs and you find every place the business rule is vague. That vagueness was going to sink the project anyway. Better to hit it in week one.

There is a commercial reason too. A buyer asking how you know the system works gets a passing suite with dates on it. That answer travels; a description of your process does not.

  • 01Label the awkward inputs, not the easy majority. A set that scores 98 on day one teaches nothing.
  • 02Keep every case the system got wrong. Those are the tests worth owning.
  • 03Set the threshold before you see the score, or it becomes whatever you just achieved.
The build orderLive
  1. CollectReal inputs from real users.
  2. AgreeA human sets the right answer.
  3. GradeBuild the scorer first.
  4. PromptNow write the instructions.
  5. GateScore drops, release stops.

Most teams run these stations in reverse and call the last two optional.

Questions, answered

Common questions

01Where do the first test cases come from if the feature does not exist yet?

From the work people are doing by hand today. Someone in the business is already answering these questions, and their past answers are your labels. Pull fifty real examples from tickets, emails or documents, and you have a starting set that reflects the actual job rather than an imagined one.

02Who decides the right answer?

A person from the business decides, not the engineering team. Whoever does the task by hand today knows which of two defensible answers the company actually stands behind. That judgement is the part no model supplies. Engineers build the grader. They should not be inventing the labels.

03How does this change the way a project is scoped?

It moves the hard conversation to the start. Agreeing what correct means for fifty real inputs takes days and it surfaces rules nobody had written down. Projects that skip it hit the same questions in week ten, with a built system and a deadline.

04Is eval-driven development worth it for a small feature?

Not always, and the honest test is what a wrong answer costs. A tone-of-voice rewrite that a person reads before sending can ship on judgement. Anything that acts without review, or that a customer sees unedited, earns a scored set before launch.

Verified
Start

Anyone can ship the agent. We answer the pager.

We build AI agents and automation, then stay on under an agreed service level. A senior engineer reads every brief, and your call gets scheduled within 24 hours.

What happens next

  1. 01

    You send a brief or book a call

    Two minutes, whichever you prefer.

  2. 02

    A senior engineer replies within 24 hours

    Not a sales rep.

  3. 03

    Honest scoping, in writing

    And if we’re not the right fit, we say so.

Abdul Basit, CEO of Hashlogics

“I started Hashlogics because too many teams ship a demo, get paid, and disappear. We build to a standard we’d run ourselves — and we stay to keep it running.”

Abdul Basit · CEO · a direct line

Not ready to talk? Take the checklist.

12 questions to ask any AI agency before you sign. They separate a demo shop from a team that ships to production.

Get the checklist

Free · no newsletter