Hashlogics
Comparison

Open-weight vs API LLMs

One question is who runs the model. This one is who owns it. Owning the weights buys freedom, not a lower bill.

The short answer

Use an API model unless you have a hard data-residency rule, GPU-scale steady volume, or a narrow task worth fine-tuning, because the break-even point sits far higher than most teams estimate.

Open weights mean you can download the model file and run it yourself. That is a real and valuable option. It is not, on its own, a cost saving, because the API price already includes serving, scaling and safety work you would otherwise take on.

Teams who compare correctly price in evaluation, GPU utilisation and the engineers who watch both. Once that is on the page, the crossover moves much further out than a spreadsheet built on GPU rental alone suggests.

Side by side

Compared on what changes the decision, not on which model tops a leaderboard this month.

DimensionAPI modelOpen-weight model
What you actually getA hosted endpoint, no fileThe weights, yours to run or move
Who runs itThe vendor, alwaysYou, or a host you pick
Where the data goesTo the vendor, under contractWherever you choose to run it
Fine-tuningLimited to what the vendor exposesFull access to the weights
Cost shapePer token, scales with useFixed hardware, plus who runs it
Licence riskOne contract, one vendorVaries by model and by version
Newest capabilityAvailable same day, most releasesTrails the frontier by months
Sensible triggerAlmost every starting pointA residency rule, or steady GPU-scale volume
Where the cost actually sitsLive
  1. WeightsThe free part. A file download
  2. ServingBatching, routing, autoscaling
  3. EvalsProof the fine-tuned model still holds
  4. GPU fleetSized for peak, idle the rest of the day
  5. On callSomeone answers when serving breaks

The weights are free. Everything after that line is the same engineering an API vendor already sells you inside the token price.

API models

Where it wins

  • You call an endpoint on day one instead of sizing a GPU fleet.
  • New model generations arrive as a version bump, not a re-deployment.
  • Serving, scaling and abuse filtering are the vendor's job, not yours.
  • Cost tracks usage, so a quiet month does not carry idle hardware.

Where it hurts

  • Your data reaches a third party, and no contract makes that fact disappear.
  • Fine-tuning is limited to whatever endpoint the vendor decides to expose.
  • A retired or changed model version can shift behaviour you tuned prompts against.
  • Cost rises with success, and one high-volume feature can dominate the bill.

Open-weight models

Where it wins

  • Data never has to leave infrastructure you control.
  • Full fine-tuning access, useful for one narrow task done very well.
  • The model stays exactly as it is until you decide to change it.
  • Steady, heavy, predictable volume can suit fixed hardware better than per-token billing.

Where it hurts

  • You are now running an inference service: sizing, serving, patching, on call.
  • Idle GPU capacity costs the same as busy GPU capacity.
  • Licence terms are not uniform across models, or even across one vendor's family.
  • The newest frontier capability usually reaches an API weeks or months before an open release matches it.

How to choose

  • Choose an API model for anything you are still proving. Buying GPUs to test an idea is an expensive way to learn the idea was wrong.
  • Choose open weights if a contract or a regulator says the data cannot leave your network, and fine-tuning has to reach the base model.
  • Choose open weights if volume is high, steady and predictable enough to keep dedicated hardware busy most of the day.
  • Choose an API model if traffic is spiky. Sizing GPUs for one peak hour and running them idle the rest of the month is a bad trade.
  • Choose both if one narrow, sensitive workload justifies fine-tuning an open model while everything else stays on an API.
  • Choose neither yet if nobody can state the accuracy the system needs. That number decides the rest of this page.
Questions, answered

Questions before you download a model

01Is an open-weight model actually cheaper?

Only at volume high and steady enough to keep dedicated hardware busy, and only once you count the engineers alongside the GPU rental. Sizing, serving, upgrades and on-call cover are ongoing work. Teams that compare on hardware cost alone consistently underestimate where the real crossover sits.

02Do open-weight models close the capability gap?

On many narrow tasks, yes: classification, extraction, structured output within a known format. On broad reasoning over messy input, the leading API models generally stay ahead by weeks or months at any given time. Test both on your own scored examples, not a public leaderboard, since leaderboard tasks rarely match yours.

03Does self-hosting an open model solve a compliance problem?

It solves the part about where data physically sits, if that is what the rule requires. It does not remove the obligation: you now own access control, logging, retention and the audit evidence a regulator asks for. Read what the rule actually demands before assuming hosting alone satisfies it.

04Can we start on an API and move to open weights later?

Yes, and that order is usually right. Keep model calls behind one internal interface from the start, so switching is a swap, not a rewrite. Build a scored evaluation set early, because without one you cannot prove the open model matches what you replaced.

05Is fine-tuning worth it if we stay on an API?

Yes, for tuning output shape or tone. Several API vendors expose fine-tuning on their hosted models, which gets you that without taking on serving. Changing the model's deeper behaviour, rather than its output format, still needs full weight access, which only an open model gives you.

Written by Abdul Basit, CEO, HashlogicsVerified
Start

Let’s build the one that runs after.

We build AI agents and automation, then stay on under an agreed service level. A senior engineer reads every brief, and your call gets scheduled within 24 hours.

What happens next

  1. 01

    You send a brief or book a call

    Two minutes, whichever you prefer.

  2. 02

    A senior engineer replies within 24 hours

    Not a sales rep.

  3. 03

    Honest scoping, in writing

    And if we’re not the right fit, we say so.

Abdul Basit, CEO of Hashlogics

“I started Hashlogics because too many teams ship a demo, get paid, and disappear. We build to a standard we’d run ourselves — and we stay to keep it running.”

Abdul Basit · CEO · a direct line

Not ready to talk? Take the checklist.

12 questions to ask any AI agency before you sign. They separate a demo shop from a team that ships to production.

Get the checklist

Free · no newsletter