Best LLM observability tools in 2026
The week after launch, someone asks why the agent said what it said. These four tools are how you answer without guessing.
The short answer
Langfuse is the strongest default for most teams, because it is open source, self-hostable, and framework neutral. LangSmith is the better pick if you already build on LangChain or LangGraph and want a managed platform that matches your stack.
Helicone wins on setup speed if the system is a single LLM call behind a proxy. Braintrust earns its place when observability and evaluation need to share one dataset.
We instrument client agents before they reach production, then keep watching after. This page covers runtime monitoring only. Pre-launch evaluation tools are ranked separately.
How this ranking was made
Verified
We ranked on what each tool captures about a live request, and how much a team can act on without building their own pipeline. Tracing nobody looks at again is not observability. So we weighted retention, search and alerting over raw feature counts.
Capabilities were read from each vendor's own documentation on 14 August 2026 and are linked in each entry. We instrument client AI systems in production, and that operational need sets the criteria below.
Generic application performance monitoring tools were excluded. They track latency and errors, not what a model retrieved or replied, and that gap is exactly the failure this page is about.
- Full trace capture
- Whether it records the input, retrieved context, tool calls and reply as one linked record, instead of only the final answer.
- Self-hosting
- Whether the traces can stay inside your own infrastructure, which matters the moment they hold customer data.
- Framework fit
- How much wiring it takes to instrument a stack that is not LangChain.
- Path to evaluation
- Whether a traced failure can become a test case in the same tool, or has to move somewhere else.
The four compared
Read from vendor documentation on 14 August 2026.
| Tool | Shape | Self-hostable | Pick it when |
|---|---|---|---|
| Langfuse | Open source platform | Yes | You want tracing you control end to end |
| LangSmith | Hosted platform | Enterprise add-on only | You already build on LangChain or LangGraph |
| Helicone | Proxy-based logging | Yes | You want logging with almost no code change |
| Braintrust | Hosted platform | Data plane only (hybrid) | Observability and evals need to share one dataset |
The ranking
Ordered by how many production teams each one suits as a starting point.
Open-source tracing you can run yourself
Start here if your stack is not LangChain, or if traces might hold data you cannot send to a third party. Langfuse's documentation describes tracing for any LLM app, plus prompt management and evaluation scoring in the same product.
Being open source and self-hostable is the real differentiator. A team in insurance or healthcare can run it inside their own network and never let a customer document leave their infrastructure. The hosted version exists too, for teams who would rather not run it.
Self-hosting is a service you now operate, and that is the trade-off. Someone has to patch it, back it up and watch its own uptime, which is a real cost even when the license is free.
Best for
- Teams that cannot send traces to a third-party platform
- Stacks built outside LangChain that still want full tracing
- Anyone who wants prompt management and evaluation in one place
Not for
- Teams with no appetite to run and patch another service
- Projects wanting zero setup on day one
- License
- Open source, self-host or hosted
- Also does
- Prompt management, evaluation scoring
The managed platform built for LangChain and LangGraph
Pick this when your agent already runs on LangChain or LangGraph. LangSmith's documentation covers tracing every run as a tree of the calls, tools and retrieval steps involved. That lines up with how those frameworks structure a request.
Instrumentation is close to automatic if you are already in that ecosystem, which is the whole appeal. It also supports online evaluation on live traffic, so a trace can be scored without you writing that pipeline.
Outside LangChain, the automatic wiring goes away. You are instrumenting by hand like any other tool, and Langfuse's open licensing becomes the stronger argument.
Best for
- Teams already built on LangChain or LangGraph
- Products wanting online evaluation on live traffic without wiring it themselves
Not for
- Stacks with no LangChain in them
- Teams that need self-hosting without an enterprise contract
- Shape
- Hosted platform
- Strongest fit
- LangChain, LangGraph
Proxy-based logging with almost no code change
Reach for this when the system is close to a single LLM call and you want logging live today. Helicone's documentation describes routing requests through its proxy, which captures the request and response without instrumenting your code by hand.
The proxy captures the model calls for free. To see a multi-step agent as one trace you add Helicone's session IDs to each step and log the tool and retrieval calls yourself, which is the instrumentation work the proxy was meant to save you.
It is the fastest of the four to turn on. That is the right trade for an early product, or a single feature that just needs someone watching it.
Best for
- A single LLM call or simple chat feature
- Teams wanting logging live in an afternoon
Not for
- Multi-step agents where you do not want to instrument tool and retrieval calls by hand
- Teams that want a full trace tree with no code changes
- Setup
- Proxy, minimal code change
- Best fit
- Single-call or simple chat features
Observability and evaluation sharing one dataset
Choose this when you want a trace you did not like to become a test case with one click. Braintrust's documentation ties logging and evaluation to the same dataset, so a production failure and a regression test live in the same place.
That link is the reason to pick it over a pure observability tool. Most teams find their best test cases in production traces anyway. Braintrust removes the export step between finding one and using it.
Self-hosting is hybrid: the data plane runs in your own AWS, GCP or Azure account, while the UI and authentication stay with Braintrust. The management plane stays with Braintrust, which is the same shape of question an enterprise LangSmith deployment raises.
Best for
- Teams who want production failures to feed evaluation directly
- Products iterating on prompts against a growing dataset
Not for
- Teams whose policy rules out any vendor-run control plane, even with data kept in their own cloud
- Anyone wanting observability with no evaluation features attached
- Shape
- Hosted platform
- Also does
- Evaluation, sharing the same dataset
- InputThe request, and who sent it.
- RetrievedPassages or records the model used.
- Tool callsWhat the agent did, in order.
- OutputThe reply, the model, the token count.
- OutcomeAccepted, edited, escalated or disputed.
Teams that log only the output can prove something happened. They cannot prove why.
When observability alone will not save you
A trace tells you what happened. It does not tell you whether that was correct, and the two questions get confused often enough to name separately.
Pairing every trace with a definition of a good answer is what evaluation does, and it belongs before launch, not after. Our roundup of evaluation tools covers that half, scored on a different set of criteria than this page.
Dashboards show the failure. Fixing the retrieval or prompt that caused it is engineering — the part our RAG development team owns on production systems.
- 01No alerting on a traced tool means someone still has to remember to look.
- 02A trace with no retention policy becomes a customer-data liability nobody decided on.
- 03Watching traces after a bad answer ships is monitoring, not prevention. Pair it with evaluation before launch.
A system where every AI decision had to be reconstructable
“I am extremely happy with the results and would highly recommend Hashlogics to anyone.”
Daniel Khin · CEO, PremiumAudit.io
Need every AI decision traceable before launch?
We wire tracing into the agent while we build it, not after a customer asks why it did something. Scoping calls cost nothing.
Questions teams ask
01Is LangSmith or Langfuse the better starting point?+
LangSmith if your agent already runs on LangChain or LangGraph and you want tracing that matches the framework with little setup. Langfuse if you want an open-source, self-hostable option that works with any stack. Data residency is usually the deciding factor.
02Do we need observability if we already run evals before launch?+
Yes. Evals check a fixed set of cases before you ship. Observability watches live traffic, where the questions and documents drift in ways no pre-launch test set predicted. They answer different questions and neither replaces the other.
03What is the minimum we should log on every request?+
The input, what was retrieved, and the final reply, linked as one record. That triplet lets you tell whether a wrong answer came from bad retrieval or bad reasoning, which is the first fork in almost every investigation.
04Can we self-host if data cannot leave our network?+
Langfuse and Helicone both support self-hosting, based on their published documentation. LangSmith self-hosts only as an Enterprise add-on, and Braintrust self-hosts its data plane in your own cloud while keeping its UI managed. Check each vendor's current terms before committing, since hosting options change.
05Why isn't Arize Phoenix on this list?+
Phoenix is a credible open-source tracing tool with its own following, particularly for teams already using its evaluation features. This ranking covers the four starting points most teams face first, split by hosting model and framework fit. Compare Phoenix against Langfuse if self-hosting is the deciding factor for you.
06Why isn't this page about evaluation tools too?+
Because they answer different questions and mixing them hides both. This page covers watching production. Our separate roundup of LLM evaluation tools covers testing before launch, which needs its own criteria.

