Find out where an AI build would stall.
Ten questions across data, evaluation, operations and ownership. No email required to see your result.
What you get
4 things that decide this
- 01A score across four areas, with the weakest one named.
- 02The specific gap most likely to stall a build, and what closing it involves.
- 03It will not tell you whether your idea is good. It tells you whether you could execute it today.
- 04Scoring well is a real result here, and it means the next step is scoping rather than groundwork.
How it scores
Four areas, weighted equally. Data covers whether the information a model would need exists, is accessible, and is correct often enough to rely on. Evaluation covers whether you could tell if a system started giving worse answers. Operations covers who watches it and who fixes it. Ownership covers whether you would hold the code, models and data at the end.
The score reports your lowest area rather than the average of the four. A build stalls at its weakest point, so an average would hide the one thing you needed to know.
The rubric is deliberately public. You could reach the same answer with a whiteboard and an honest hour, and a score you cannot reproduce is a lead-capture trick rather than an assessment.
- DataExists, accessible, correct enough
- EvaluationYou would notice it getting worse
- OperationsNamed owner, agreed service level
- OwnershipCode, models and data are yours
Equal weight, lowest score reported. That is where a build actually stops.
What a strong and a weak answer look like
Published so you can score yourself. Two or three weak answers in one area is the result worth acting on.
| What is asked | Strong answer | Weak answer |
|---|---|---|
| Where the data lives | One system of record, with a named owner | Several systems that disagree, reconciled by hand |
| How wrong the data is | A known error rate, checked on a schedule | Nobody has measured it |
| How you would catch a bad answer | Real cases with known-correct answers, re-run on every change | Someone would spot-check it if a customer complained |
| How a failure reaches you | Alerts on error rate and latency, not only downtime | A user tells you |
| Who operates it | A named person and an agreed service level | The team that built it, informally |
| What you hold at the end | Code, prompts, models and pipelines yours in writing | Access to a hosted product and nothing underneath |
Walk through it with an engineer
Whatever the score says, the call is free and you keep the assessment.
Common questions
01What happens if we score well?
Then you are in good shape, and the next step is scoping the build rather than fixing foundations. Scoring everyone as unready would make this a sales trick. We would rather tell you that you are ready and be right.
02Do you benchmark our score against other companies?
No. We have no dataset that would make a comparison honest, and an invented benchmark is the kind of claim that cannot be defended. The score is measured against what a production build requires, not against other people.
03Which area do most organisations score worst on?
Evaluation, in our experience of scoping calls. Teams often have workable data and no way to tell whether a system is still correct next month. That gap is exactly what turns a working demo into a quiet failure.
04Can we score ourselves without talking to anyone?
Yes, and the rubric above is published for that reason. Work through the six rows with the people who own the data and the on-call rota, and be honest about the weak answers. A score you talk yourself into is worth nothing.

