Hashlogics
Blog

De-identification is where clinical AI projects get real

The roadmap survives the pitch meeting. The first hard engineering question shows up when someone asks what the model actually sees.

In short

5 things that decide this

  1. 01Every clinical AI plan meets the same question at the same point: what exactly does the model see?
  2. 02HIPAA names two ways to answer it. Safe Harbor strips 18 kinds of identifier. Expert determination has a qualified statistician certify that re-identification risk is very small.
  3. 03The cheaper route for most pipelines is Safe Harbor. It also removes information a model could use, like a precise age or a full ZIP code.
  4. 04TrialTriage, built by Hashlogics, holds patient data as age ranges, ZIP prefixes and pseudo-patient IDs, then adds field-level encryption and a nurse's final review on top.
  5. 05Re-identification risk does not end at the de-identification step. A pipeline that logs prompts or keeps free-text notes can reintroduce identifiers a table never had.
The setup

The question every plan meets eventually

A clinical AI pitch usually gets through the first meeting on the strength of the outcome. Faster trial matching, faster chart review, fewer hours spent cross-referencing guidelines by hand. Nobody objects to that part.

The project gets real at a narrower question. What exactly reaches the model? A name and a birth date are the obvious answer, and most teams strip those on instinct. HIPAA's own rule, at 45 CFR 164.514, goes further than instinct does. That gap is where a plan turns into an architecture.

The mechanism

Two routes, and they are not interchangeable

HHS guidance under the Privacy Rule sets out two ways to de-identify health data. A build has to pick one on purpose, not drift into a mix of both.

The first is Safe Harbor. It removes 18 named types of identifier: name, exact dates, full ZIP code, phone number, email, and more down to biometrics and full-face photos. What is left must carry no way to trace back to a person. It is a checklist, and a checklist is easy to audit.

The second is expert determination. A trained statistician checks the data and signs off that the risk is very small. This route can keep more detail, a wider age range or a fuller ZIP code. It also costs an expert's time and a paper trail that has to hold up on review. Most clinical AI teams pick Safe Harbor instead. It needs no specialist to sign off on every dataset.

  • 01Safe Harbor: 18 types of identifier stripped out, checked off a list
  • 02Expert determination: a trained person signs off that the leftover risk is very small
  • 03The cheaper default for a pipeline that runs all the time is Safe Harbor
  • 04Neither method survives free text left unscrubbed. Clinicians write dates and names into notes out of habit
The evidence

What this looked like in TrialTriage

TrialTriage matches oncology patients to clinical trials against NCCN guidelines and drug data. Nurses used to do that search by hand, and the delay risked a patient missing an enrollment window.

Hashlogics built the data model around Safe Harbor from the start. Age becomes a range instead of a birthdate. Location becomes a ZIP prefix instead of a full ZIP code. A pseudo-patient ID replaces the name everywhere the model or an insurer's batch job touches the record. The whole store sits behind field-level encryption. A nurse still reviews and finalizes every ranked match before it reaches a patient or an insurer.

None of that was added later. It shaped the schema before the matching code was written. That is why the audit trail tracked 23 kinds of action from day one, instead of arriving after a security review asked for one.

The fix

Where teams reintroduce the risk they just removed

Stripping the table is not the end of the job. Prompt logs, tracing tools and dashboards copy whatever text goes to the model. That text can carry identifiers the table never had, especially once a free-text note joins the mix. Scrub where a call leaves your system, not inside each developer's habits.

Pick the method before the schema is built, not after a review asks which one you used. Safe Harbor for most pipelines. Expert determination only where the extra detail earns the extra process. Either way, a licensed person still signs off on what the model produces. Stripping the data does not make the output correct.

Questions, answered

Questions this raises

01What is the difference between Safe Harbor and expert determination?

Safe Harbor removes 18 named types of identifier, like names, exact dates and full ZIP codes, with no known way left to trace the data back to a person. Expert determination works differently: a trained statistician signs off that the leftover risk is very small. That route can keep more detail, but it needs a documented expert review for each dataset.

02Does de-identified data take a project out of HIPAA?

Yes, if it is de-identified by one of those two methods, and nobody involved knows a way to link it back to a person. A dataset a team knows can be re-identified stays in scope, no matter which columns were dropped.

03Why does an LLM pipeline need de-identification if the model provider signs a BAA?

A business associate agreement covers how the provider handles data it receives. It does not reduce what the model actually sees. Stripping identifiers before the call shrinks the damage a provider incident could do. It also cuts what a prompt log or a cached response could ever expose.

Written by Abdul Basit, CEO, HashlogicsVerified
Start

Let’s build the one that runs after.

We build AI agents and automation, then stay on under an agreed service level. A senior engineer reads every brief, and your call gets scheduled within 24 hours.

What happens next

  1. 01

    You send a brief or book a call

    Two minutes, whichever you prefer.

  2. 02

    A senior engineer replies within 24 hours

    Not a sales rep.

  3. 03

    Honest scoping, in writing

    And if we’re not the right fit, we say so.

Abdul Basit, CEO of Hashlogics

“I started Hashlogics because too many teams ship a demo, get paid, and disappear. We build to a standard we’d run ourselves — and we stay to keep it running.”

Abdul Basit · CEO · a direct line

Not ready to talk? Take the checklist.

12 questions to ask any AI agency before you sign. They separate a demo shop from a team that ships to production.

Get the checklist

Free · no newsletter