The best LLM observability tools in 2026
Seven tools, with entry prices running from free to $249 a month, and none of them will fix a bad prompt for you. What they do is show you which prompt, which model and which tool call produced the bad output, and what it cost you to get it.
This list is for anyone with an LLM call inside something that matters: a support bot, an agent that touches customer records, a batch job that rewrites pages overnight. If your debugging method is print statements and “it worked when I tried it”, you’re the reader. By LLM observability I mean tracing each request (prompt, model, parameters, nested tool calls, tokens, latency, cost) and then attaching a score or a human verdict afterwards.
I’m Xavier, an operator in Singapore, and I run batch-style LLM pipelines. That biases me toward cost per run, prompt versions and finding the one bad output in a big batch, which is a different job from watching a live chat product. One limit up front: this comparison rests on vendor docs, pricing pages and what each tool asks of your code. I haven’t load-tested any of them or measured overhead. Prices below are list prices and this category reprices often, so confirm on the vendor page. More writing like this is on the blog index.
how I picked
Six criteria, roughly in the order I’d weigh them for a small team.
- tracing depth: nested spans for agent loops, tool calls and retrieval steps, not a flat list of requests
- evals: scores, an LLM judge or a human review queue, and a way to compare two prompt versions on the same dataset
- exit cost: whether it ingests OpenTelemetry. The OpenTelemetry GenAI semantic conventions define standard span attributes for model calls and were still evolving when I last read them, so expect attribute names to shift. Tools that speak them are easier to leave later
- hosting: a self-host option or a regional cloud, because prompts tend to contain customer data
- price shape: what a small team pays versus what a busy batch job pays, since those are different bills
- setup time: an afternoon, not a quarter
the picks
Langfuse
Langfuse is where I’d send most small teams first. The core is open source under an MIT license, so you can run it yourself with Docker Compose or Helm, or use the hosted cloud. It covers nested tracing, prompt versioning with a playground, datasets and experiments, LLM-as-judge evaluators and user feedback, with Python and JS/TS SDKs, an OpenTelemetry endpoint and integrations for the OpenAI SDK, LangChain, LlamaIndex and LiteLLM.
Prompt management is the feature that matters for pipeline work. You fetch prompts at runtime, edit them in the UI, and every trace records which version produced it. When quality drops on a Tuesday, you can check whether someone edited the prompt on Monday.
- pro: open source core with a real self-hosting path, useful when prompts hold customer data
- pro: prompt versions, traces and scores in one place
- pro: broad integrations plus OpenTelemetry ingest
- con: self-hosting the current version means running ClickHouse, Postgres, Redis and object storage, which is real ops work
- con: billing counts traces, observations and scores as units, so forecasting takes arithmetic
pricing: Hobby is free with 50k units a month, Core is $29 a month, Pro is $199 a month and Enterprise starts at $2,499 a month, per the Langfuse pricing page. Self-hosting the open source version costs only your own infrastructure.
read next: how to write a prompt you can maintain
LangSmith
LangSmith is LangChain’s platform, and it works without LangChain too: a traceable decorator, wrappers for the OpenAI and Anthropic clients, and OpenTelemetry support. It’s at its best with LangChain or LangGraph, where setting LANGSMITH_TRACING=true and an API key gets you traces with no code changes, and agent runs render as a graph with state at each step.
Beyond tracing you get datasets, offline experiments, online scoring on live traffic, a prompt hub and annotation queues for human review. That last one is why I’d shortlist it when the people scoring outputs aren’t engineers.
- pro: two environment variables give you traces from a LangChain or LangGraph app
- pro: annotation queues make human review a normal workflow
- pro: offline experiments and online scoring on live traffic in the same tool
- con: self-hosting is an Enterprise feature, so teams with strict data rules end up on a sales call
- con: per-seat plus per-trace pricing grows with every reviewer you add
pricing: Developer is free for one seat with a monthly allowance of 5k base traces, Plus is $39 per seat a month and Enterprise is custom. Extra traces are billed per thousand, at a higher rate for longer retention. Check LangSmith’s pricing page for current numbers.
read next: comparing agent frameworks by what they lock you into, because the framework you pick decides how easy tracing is
Arize Phoenix
Phoenix is Arize AI’s open source tool, built on OpenTelemetry and Arize’s OpenInference instrumentation conventions. It runs from a notebook, a Docker container or the hosted Phoenix Cloud, and the local version needs no account. You get tracing, LLM-judge eval templates, datasets, experiments and a prompt playground.
I like it for the debugging stage, especially retrieval. You can see which chunks came back for a query and score whether they were relevant, which is where a lot of RAG failures hide. Arize AX is the commercial platform above it, with production monitoring, and it’s a separate product with its own bill.
- pro: OpenTelemetry-based instrumentation, so traces aren’t trapped in one vendor’s SDK
- pro: runs on a laptop with no account, good for debugging a pipeline before you deploy it
- pro: retrieval and RAG inspection is a genuine strength
- con: the Elastic License 2.0 is source-available rather than OSI open source, and it restricts offering Phoenix as a hosted service
- con: production monitoring means moving to Arize AX, so cost planning happens in two steps
pricing: Phoenix is free to self-host and Phoenix Cloud has a free tier. Arize AX has a free plan, AX Pro is listed at $50 a month and Enterprise is custom.
read next: how to chunk documents for retrieval
Braintrust
Braintrust is evals-first. Production logs feed datasets, datasets feed experiments, and experiments compare prompt or model changes with scorer output side by side. There’s a GitHub Action to run evals in CI, so a prompt edit that drops your score can fail the build. The scoring library, Autoevals, is open source.
If your team’s recurring question is “did that prompt change help”, this is the one to try. Its tracing is fine, but I’d pick something else if all you need is a window into production.
- pro: experiments show side by side diffs, which turns “is this prompt better” into a concrete question
- pro: production logs become eval datasets without much ceremony
- pro: the free tier includes unlimited users, so reviewers don’t need seats
- con: closed source, and running it in your own cloud is an Enterprise arrangement
- con: the step from free to $249 a month is flat and steep, and the data caps are worth checking if your prompts are long
pricing: Starter is free with 1 GB of processed data and 10k scores, Pro is $249 a month with 5 GB and 50k scores, and Enterprise is custom. Usage beyond the included amounts is charged extra.
read next: when to put a human in the loop
Helicone
Helicone is the quickest thing on this list to try. In proxy mode you change the base URL your OpenAI-style client points at, add an auth header, and requests start showing up with cost, latency and per-user breakdowns. It also does caching, rate limits and retries at the gateway, and groups multi-step runs into sessions. The project is open source under Apache 2.0 and can be self-hosted.
It’s what I’d install if someone asks “what is this costing us and who is spending it” and wants an answer today. The tradeoff is that proxy mode puts a third party in your request path. An async logging option keeps it out of the path, but that takes more wiring.
- pro: setup is often a one-line base URL change
- pro: cost by user, model and prompt from day one, and caching trims repeat spend
- pro: open source with a self-host path
- con: evals and dataset tooling are thinner than Langfuse or Braintrust
- con: proxy mode adds a hop and a dependency to every request
pricing: Hobby is free at 10k requests a month, Pro is $79 a month, Team is $799 a month and Enterprise is custom.
read next: when a small model is the right choice, since per-request cost data is what tells you whether a cheaper model would do
Datadog LLM Observability
Datadog LLM Observability makes sense for one kind of buyer: a team already paying for Datadog. An LLM call shows up as a span in the same trace as your web service, database and queue, so a slow response can be traced to the model, the retrieval step or the SQL query behind it. Auto-instrumentation through the ddtrace library covers common providers and frameworks, and there are built-in quality and safety checks for things like hallucination, prompt injection and sensitive data leaks. The details are on the Datadog LLM Observability page.
If you aren’t already on Datadog, I wouldn’t adopt it just for this. Billing is per monitored LLM request on top of your existing plan, which gets expensive at batch volumes, and the workflow feels aimed at operating a live service rather than iterating on prompts.
- pro: LLM spans sit in the same trace as services, databases and infrastructure metrics
- pro: alerting, dashboards and on-call routing you already have
- pro: built-in quality and safety checks
- con: per-request billing on top of an existing Datadog plan hurts at volume
- con: prompt iteration is a secondary workflow, so you may still want a second tool
pricing: billed per monitored LLM request, with rates that depend on term and volume. I won’t quote a number I can’t verify, so model your request volume on Datadog’s pricing page and ask for a quote.
read next: when a bigger window replaces your retrieval layer, because long prompts show up in traces as token cost and latency
Weights & Biases Weave
Weave is Weights & Biases’ tracing and evaluation toolkit for LLM apps. Call weave.init with a project name and it patches supported libraries such as OpenAI, Anthropic and LangChain, then you decorate your own functions with @weave.op to see them in the trace. Evaluations use datasets and scorers, with comparison views across runs. There are Python and TypeScript SDKs, and the client library is open source under Apache 2.0, though hosted W&B is the main way to use it.
It makes the most sense if your team already tracks training or fine-tuning runs in W&B, so LLM traces sit next to model experiments. W&B is now owned by CoreWeave, so ask about roadmap and pricing before you sign anything annual.
- pro: one init call and a decorator get you traces
- pro: traces and evals sit beside training and fine-tuning runs if you already use W&B
- pro: Python and TypeScript SDKs
- con: with no other reason to use W&B, it’s another account and another vocabulary
- con: pricing follows W&B’s plan structure and data ingested, which is harder to map to LLM traffic than a per-trace price
pricing: there’s a free tier for personal use. Paid plans depend on the plan and how much data you ingest, so check W&B’s pricing page. I’m not quoting a number.
read next: ai benchmarks and why they mislead, because your own eval dataset tells you more than a leaderboard
comparison table
| tool | entry price | primary strength | primary weakness |
|---|---|---|---|
| Langfuse | free, then $29/month | open source with prompt management | self-hosting needs several services |
| LangSmith | free, then $39/seat/month | LangChain and LangGraph tracing | self-hosting only on Enterprise |
| Arize Phoenix | free to self-host | OpenTelemetry based, retrieval debugging | source-available license, separate paid AX |
| Braintrust | free, then $249/month | evals and experiments | closed source, steep flat price step |
| Helicone | free, then $79/month | fast setup and cost tracking | thin evals |
| Datadog LLM Observability | per request, on top of Datadog | one trace across your whole stack | cost at volume |
| Weights & Biases Weave | free personal tier | sits beside W&B experiment tracking | little value if you don’t use W&B |
how to choose
Start with who reads the traces. If it’s only engineers chasing bugs, Langfuse, Phoenix or Helicone are enough. If product managers or domain reviewers need to score outputs, look hard at LangSmith’s annotation queues and Braintrust’s free unlimited users.
Then decide where the data can live. Prompts contain customer names, order details, sometimes worse. From Singapore I’d check where each vendor’s cloud region sits before signing anything, and if the answer isn’t good enough, self-hosting Langfuse, Phoenix or Helicone keeps traces on your own servers for the price of running them. This is not legal advice, and a PDPA question deserves a proper answer from someone qualified. For the wider privacy angle, our sister site The Privacy Wire is the place to look.
Keep your exit cheap. Wrap your model calls in one small function of your own so tracing lives in a single place, and prefer tools that ingest OpenTelemetry. This is a young and consolidating category, so pick monthly billing over an annual contract until you’ve seen a release cadence you trust. Read the repo’s commit history and the changelog, whatever the logo on the website says.
Last, look at volume, because trace count drives most of these bills. Sample successful runs, keep every error and every low-scored run, and shorten retention. I’d also argue against running an LLM judge on every trace. It adds a second model call to each one, and you shouldn’t trust the score until a person has compared it with a sample. Put two candidates on their free tiers against a real week of your traffic and compare the bills you’d get.
verdict / top pick
Langfuse is my top pick for most small teams. It has the best balance here of open source core, prompt versioning, evals and a cheap entry price, $29 a month for cloud, and if data rules bite you can host it yourself. The self-hosting stack is heavier than it looks, so budget for that.
The exceptions are fairly clear. Live in LangChain or LangGraph with non-engineer reviewers, and LangSmith wins. Need to prove that prompt changes help, and Braintrust is built for it. Already paying Datadog, add its LLM product before adding a new vendor. Helicone is for cost answers today, Phoenix for retrieval debugging on a laptop, and Weave if W&B is already open in another tab.
Written by Xavier Fok
disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-09-26.