Hashlogics
Hire

LLM engineering

Prompting is the part that takes an afternoon

Our engineers run large language models against ESG research, insurance documents and a live search retrofit. The work is making the output measurable, affordable and stable when the model underneath changes.

What you are getting

4 things that decide this

  1. 01Senior engineers with language models in production across research, document and search products described on this site.
  2. 02They write the scoring set before they tune the prompt, because a change you cannot measure is a change you cannot defend.
  3. 03You interview each of them yourself, and you may decline without giving a reason.
  4. 04Prompts, scoring sets and pipelines sit in your repository from the first commit.

What an LLM engineer actually does here

A prompt is the cheapest part of the system and the least stable. Everything valuable is built around it.

On Greenlight the platform researches more than 50 ESG topics, pulling 10 to 15 independent sources for each so a score never rests on a company describing itself. On PremiumAudit, Claude reads insurance audit documents where a misread number has a consequence. On Military Cruise Deals, conversational search was fitted into a running WordPress site using FastAPI and LangGraph.

The repeated skill is not phrasing. It is deciding what evidence the model is allowed to use, how the answer gets checked, and what the system does when confidence is low.

Where the real work sits

Scoring before tuning

A set of real cases with expected outcomes. Prompt changes then get accepted or rejected on evidence instead of on whoever spoke last.

Output that fits a schema

Structured output validated before use. A free-text answer parsed with string matching breaks the first time the model phrases it differently.

Sources the answer must rest on

Evidence gathered and cited, as on Greenlight, so a claim can be traced. Self-reported data alone is how a research tool becomes a marketing tool.

Spend that does not surprise you

Smaller models on easy work, caching for repeats, and a ceiling on what one request may cost. Usage-based bills grow quietly.

Version changes handled on purpose

Model versions are retired on a published schedule. Running the scoring set against the replacement turns that into a decision you make on a calm day.

How an answer gets accepted or rejectedLive
  1. QuestionFrom a user or a job
  2. EvidenceWhat it may rely on
  3. Model callStructured output asked for
  4. ValidateShape and facts checked
  5. ScoreAgainst known cases
  6. Ship or refuseLow confidence goes to a person

Boxes four and five are what most teams add after their first bad week in production. Building them first is cheaper.

A client, in their own words

I am extremely happy with the results and would highly recommend Hashlogics to anyone.

Daniel Khin · CEO, PremiumAudit.io

How hiring works

  1. 01

    Show us what it gets wrong

    A free call about the outputs that embarrass you and the ones you cannot check. If the fix is your data rather than your prompt, we say that.

  2. 02

    Meet the engineers

    We shortlist people who have measured model quality against real cases, and you interview them your own way.

  3. 03

    They embed

    Your repository, your standups, your review process. One of our engineers owns the scoring set and what it says.

  4. 04

    They hand over

    The scoring set, the prompts under version control and a note on what each model version costs you. Where a client wants the quality watched after launch, we do that under a service level we agree.

What they work with

Stack

Models

ClaudeOpenAI GPT-4oFireworks AIPerplexity Sonar-Pro

Orchestration

LangGraphFastAPIPythonRedisPostgreSQL

Practices

Scored case setsStructured outputPrompt version controlSpend ceilings
Next step

Bring us the output you cannot trust yet

Send the answers that are wrong and the ones you have no way to check. The scoping call is free, and you will leave knowing whether this is a prompt problem or an evidence problem.

Questions, answered
01Is an LLM engineer different from an AI engineer?

The titles overlap and the difference is emphasis. An LLM engineer spends their time on model behaviour: prompting, structured output, evaluation, cost and version changes. An AI engineer covers a wider system including retrieval and the surrounding product. We staff both and will tell you which one your problem actually needs.

02Should we fine-tune a model or improve the prompt and data?

Improve the data first, nearly always. Fine-tuning is worth it when you need a consistent format or a narrow style at scale and you have clean examples to train on. It does not add knowledge, which is the mistake behind most fine-tuning requests we are asked to price.

03How do you keep our data out of model training?

By choosing the API terms deliberately and checking them rather than assuming. Both Anthropic and OpenAI state that API data is not used to train their models unless you opt in. That is the default for the API, and consumer chat products are governed separately, which is where teams get caught. We read the current terms for your provider during scoping and write the answer into the design.

04Can you work with an open-weight model we host ourselves?

Yes. It suits teams with strict data rules or heavy predictable volume where per-token pricing hurts. The trade is that hosting, scaling and upgrades become yours, and quality on hard reasoning tasks may sit behind the strongest hosted models. We compare both against your actual cases rather than arguing in the abstract.

05How do you control what we spend per month?

Route easy work to smaller models, cache repeated calls, cap the tokens one request may use, and track spend per feature rather than as one bill. Greenlight pulls many sources per topic, so the routing decision there is what keeps a research run affordable.

06What sets the price of this work?

How measurable your task is. A job with a clear right answer can be scored quickly and tuned with confidence. A judgement task where two experts disagree needs more work up front to define what good even means. Scoping calls are free. Where the answer needs us inside an existing codebase, a paid two-week diagnostic replaces the estimate with a fixed price.

Verified
Start

Anyone can ship the agent. We answer the pager.

We build AI agents and automation, then stay on under an agreed service level. A senior engineer reads every brief, and your call gets scheduled within 24 hours.

What happens next

  1. 01

    You send a brief or book a call

    Two minutes, whichever you prefer.

  2. 02

    A senior engineer replies within 24 hours

    Not a sales rep.

  3. 03

    Honest scoping, in writing

    And if we’re not the right fit, we say so.

Abdul Basit, CEO of Hashlogics

“I started Hashlogics because too many teams ship a demo, get paid, and disappear. We build to a standard we’d run ourselves — and we stay to keep it running.”

Abdul Basit · CEO · a direct line

Not ready to talk? Take the checklist.

12 questions to ask any AI agency before you sign. They separate a demo shop from a team that ships to production.

Get the checklist

Free · no newsletter