Hashlogics
Computer vision development

Vision systems judged on your worst images

You get a system scored against the pictures your cameras really take, not a benchmark. Where a mistake costs real money, a person approves before anything happens.

The standard

Accuracy on a slide is not accuracy in a warehouse. A model demoed on clean, well-lit, straight-on images meets something else in production: glare, a smudged lens, a form printed at an angle on a bad photocopier. Your real number is the one measured on those. So the first thing we build is a set of your own images, labelled by your own people. The passing bar gets agreed before anyone sees a result. Everything after that is engineering.

The problem

The percentage that means nothing on its own

Vendors quote one accuracy number. It hides the thing that decides whether the system is usable: which kind of mistake it makes.

A system that occasionally flags a good part for review costs you a few seconds of a person's time. A system that occasionally passes a bad one costs you a recall. Same headline number, completely different business.

So the useful conversation is about what the system should do when it is unsure. A vision system allowed to say I do not know is worth more than a confident one that is right slightly more often.

What we have shipped

Counted, not estimated

22

systems in production

9

of them with AI doing real work inside the product

1

insurance document pipeline reading audit paperwork: PremiumAudit.io

23

tracked actions in TrialTriage's audit trail

The work

Where vision earns its place

The problems where a camera beats a person, and the ones where it does not.

Reading documents that arrive as pictures

Scans, photos of forms, PDFs that are really images. The work is not the reading. It is the field mapping, the validation, and what happens to the ones that come out wrong.

Inspection and counting

Checking a part, a shelf or a site against what should be there. These pay off where the count happens often and a person doing it is expensive or slow.

Detecting a condition, not identifying a person

Spotting a spill, an empty bay, a missing safety item. We steer clients away from identifying individuals, because the legal exposure usually outweighs what it buys.

Knowing when to stop and ask

Every build gets a confidence threshold and a queue. Below the line, a person looks. That queue is also where your next batch of training data comes from.

How a vision build runsLive
  1. CollectYour images, including the bad ones.
  2. LabelYour experts decide what correct means.
  3. ScoreBoth error types, counted separately.
  4. GateUnsure cases go to a person.
  5. WatchCameras move. Accuracy drifts.

Score counts both error types separately, because one number hides the difference between flagging a good item and passing a bad one. Those have very different costs.

The hardest part

The day someone moves the camera

Vision systems degrade in ways software does not. Nobody changed the code. Someone changed a light fitting, or cleaned a lens, or replaced the forms with a new layout.

Accuracy falls and nothing raises an error, because the system is doing exactly what it was built to do on inputs it has never seen. That is why the review queue matters after launch as much as during the build. When the share of unsure cases climbs, something in the physical world moved. You find out in days, not at the next audit.

  • The share of low-confidence cases is tracked as a health signal, not just a workload.
  • New camera positions and new form layouts get re-scored before they go live.
  • Corrections from the review queue feed the next version.
  • Where images contain personal data, we keep what the model does not need out of reach.
The stack

What we build on

Models

Claude API visionOpenAI visionOCR pipelinesOpen-source detection models

Build

PythonFastAPIPyMuPDFCelery + RedisPostgreSQL

Run

AWSAWS S3DockerSentryGitLab CI

Around it

Review queuesConfidence thresholdsScored test setsAudit logging
A client, in their own words

I am extremely happy with the results and would highly recommend Hashlogics to anyone.

Daniel Khin · CEO, PremiumAudit.io

The usual vision pitch against ours

Both show a working demo. They separate on how the number was produced.

CriterionThe usual approachHow we build
The accuracy figureFrom a public benchmark or the vendor's own images.From your images, labelled by your people, agreed in advance.
Error typesOne number, both kinds averaged together.Counted separately, because their costs are not the same.
When unsureGuesses, and records a result.Sends it to a person, and logs why.
When the camera movesFound out at the next audit.The unsure rate climbs and raises an alert within days.
Faces and peopleOffered, because it sells.Argued against, unless there is a lawful reason we can point to.
Questions, answered
01Have you shipped a dedicated computer vision product?

Our vision work sits inside document-heavy AI systems rather than as a standalone product line, and pretending otherwise would be easy to check. PremiumAudit.io reads insurance audit paperwork with field mapping, validation and exception handling on Claude, with a human auditor in the loop. That is the shape of vision problem we have production experience with. Ask for something else and we will tell you where the edge of our experience is.

02How many images do you need before we start?

Fewer than most people expect for a first assessment, and more than they expect for a reliable one. A few hundred labelled examples covering your genuinely hard cases will tell you whether the problem is solvable at all. The messy examples matter far more than the clean ones, because a model gets the clean ones right anyway.

03Do we need to train our own model?

Usually not any more. General-purpose vision models handle a lot of document and scene work with no training at all. That turns a six-month project into an evaluation. Training your own is worth it when the objects are specific to your business and no general model has seen anything like them.

04What accuracy can we expect?

Nobody can answer that honestly before seeing your images, and a vendor who quotes a number up front is quoting someone else's data. The method is what we commit to. Score on your examples, count both error types separately, and agree the bar before results are visible. If the score does not clear it, that is the finding.

05Can this run on cameras we already have?

Often yes, and it is worth checking before buying anything. Existing cameras fail mainly on resolution, on lighting, or on an angle that hides what matters. Those are cheap to test with a sample of real footage, and that test costs far less than a hardware refresh.

06What about privacy when the images contain people?

Keep the personal data away from the parts of the system that do not need it, and be clear about what is kept. We build to what the standard asks for, and the system emits the evidence an auditor looks for. Where a workflow only needs to know that a bay is empty, it should never receive a face at all.

07How does the system improve after launch?

Through the review queue, which is why every build has one. Cases the system was unsure about get corrected by a person, and those corrections become the next round of test and training data. A vision system with no queue has no route to getting better, and no way to notice it is getting worse.

08Is this only worth it at high volume?

Volume matters, but the cost of the mistake matters more. A hundred documents a day where an error triggers a costly rework can justify the build faster than ten thousand low-stakes images. On the free scoping call we work out which side of that line you sit on, and we will say when the answer is neither.

Verified
Start

Anyone can ship the agent. We answer the pager.

We build AI agents and automation, then stay on under an agreed service level. A senior engineer reads every brief, and your call gets scheduled within 24 hours.

What happens next

  1. 01

    You send a brief or book a call

    Two minutes, whichever you prefer.

  2. 02

    A senior engineer replies within 24 hours

    Not a sales rep.

  3. 03

    Honest scoping, in writing

    And if we’re not the right fit, we say so.

Abdul Basit, CEO of Hashlogics

“I started Hashlogics because too many teams ship a demo, get paid, and disappear. We build to a standard we’d run ourselves — and we stay to keep it running.”

Abdul Basit · CEO · a direct line

Not ready to talk? Take the checklist.

12 questions to ask any AI agency before you sign. They separate a demo shop from a team that ships to production.

Get the checklist

Free · no newsletter