AI Monitoring and Observability

Monitoring Tools and Platforms Overview


A twelve-person team ships an LLM product. Their month-three observability bill is 2,800 dollars — twenty hosts of infrastructure and APM monitoring, plus logs. Reasonable. Month four the bill is 9,100 dollars. Traffic grew 6%.

The cause is one line in a pull request. An engineer wanted per-customer latency graphs, so they added a customer_id tag to four metrics. Those metrics already carried a model label (5 values) and an endpoint label (8 values), so they represented 4 × 5 × 8 = 160 time series. The company has 800 customers:

Text
160 series x 800 customers          = 128,000 custom metricsincluded with 20 hosts (100 each)   =   2,000billable                            = 126,000at 0.05/metric/month                =   6,300/month

Six thousand three hundred dollars a month for a graph that was viewed twice. The same information was already sitting in their logs, sliceable for free.

Choosing observability tools is not really a feature comparison. Every product in this space can draw a line chart. What differs is the cost model, the data model, and what happens when you want to leave — and those three things determine whether the stack you pick survives contact with growth.

Four jobs no single tool does wellYourobservability stackMetrics —Prometheus, pull modelDashboards — GrafanaLogs — indexall, or index noneLLM traces —prompts, tokens, costOpenTelemetry —the escape hatch
The bill went from 2,800 to 9,100 dollars on 6 percent more traffic because these tools charge for cardinality, not requests.

Four jobs, and why no single tool does all of them

People say "we use Datadog" or "we use Prometheus" as though it were one decision. It is four, and the four have genuinely different storage requirements.

JobData shapeQuery patternScales withTypical tools
Metrics storeNumeric time series, fixed labelsAggregate over time, group by labelActive series (cardinality)Prometheus, Mimir, VictoriaMetrics, Datadog
Log storeSemi-structured events, arbitrary fieldsFilter, full-text search, countBytes ingested and indexedElasticsearch/OpenSearch, Loki, Datadog Logs, Splunk
Trace storeSpan trees, high cardinality by natureFetch one trace; aggregate over span attributesSpans stored (after sampling)Jaeger, Tempo, Datadog APM, New Relic
Model / LLM layerPrompts, completions, scores, datasets, drift statsCompare runs, inspect examples, score evalsRequests logged, eval runsLangSmith, Langfuse, Arize Phoenix, Evidently

A metrics database that also indexes arbitrary text is a search engine wearing a costume, and it will be bad at one of the two. The reason "one pane of glass" products exist is that switching context during an incident is genuinely expensive — but internally they are still running four different storage engines, and you pay for each.

Prometheus: the pull model and what it buys you

Prometheus is the default open-source metrics store, and its most distinctive design choice is that it pulls. Your application exposes a plain-text endpoint at /metrics; Prometheus scrapes it on a schedule.

Text
# HELP llm_requests_total Total LLM requests# TYPE llm_requests_total counterllm_requests_total{model="assistant-v4",endpoint="/chat",status="ok"} 148203llm_requests_total{model="assistant-v4",endpoint="/chat",status="error"} 71# HELP llm_latency_seconds End-to-end latency# TYPE llm_latency_seconds histogramllm_latency_seconds_bucket{model="assistant-v4",le="0.5"} 91442llm_latency_seconds_bucket{model="assistant-v4",le="1.0"} 132890llm_latency_seconds_bucket{model="assistant-v4",le="+Inf"} 148274llm_latency_seconds_sum{model="assistant-v4"} 96331.4llm_latency_seconds_count{model="assistant-v4"} 148274

Pull sounds like an implementation detail and is not. Four consequences:

  • Scrape failure is itself a signal. If a target stops responding, Prometheus records up = 0. With push, a silent instance is indistinguishable from a healthy quiet one.
  • Back-pressure is automatic. An overloaded Prometheus scrapes more slowly. It does not get flooded by applications pushing harder.
  • You can run the endpoint by hand. curl localhost:8000/metrics during development shows exactly what will be collected, which makes instrumentation bugs obvious in seconds.
  • Short-lived jobs do not fit. A batch job that runs for 40 seconds may never be scraped. That is what the Pushgateway is for — and it is a genuine exception, not a general-purpose push endpoint. Metrics pushed there persist until deleted, so using it for ordinary services gives you stale values forever after a deploy.

The architecture is small: a scrape loop feeding a local time-series database, a rule evaluator computing recording and alerting rules on a schedule, an HTTP query API, and Alertmanager as a separate process handling grouping, silencing, and routing.

Retention arithmetic, so you can size it

Prometheus compresses aggressively — roughly 1.5 bytes per sample after delta-of-delta and XOR encoding. With 100,000 active series and a 15-second scrape interval:

Text
samples/sec  = 100,000 / 15                = 6,667samples/day  = 6,667 x 86,400              = 576,000,000bytes/day    = 576,000,000 x 1.5           = 864 MB/day30-day disk  = 864 MB x 30                 = 25.9 GB

Twenty-six gigabytes for a month of a hundred thousand series. This is why a single Prometheus on a modest box goes a very long way, and why the framing "we need a managed metrics platform" often means "we have a cardinality problem" rather than "we have a scale problem".

Watch what the earlier tagging mistake does to the same calculation at 12 million series:

Text
samples/sec  = 12,000,000 / 15             = 800,000samples/day  = 800,000 x 86,400            = 69,120,000,000bytes/day    = 69.12e9 x 1.5               = 103.7 GB/dayRAM for the active series head block (~3.5 KB each) = ~42 GB

The disk is survivable. The 42 GB of resident memory is not, and this is how Prometheus servers die: not gradually, but by out-of-memory kill during a deploy that adds one label.

Prometheus scales beautifully with time and traffic and catastrophically with cardinality. Every capacity conversation about a metrics store is really a conversation about label values.

PromQL, and recording rules

SQL
-- error ratio over 5 minutessum(rate(llm_requests_total{status="error"}[5m]))  / sum(rate(llm_requests_total[5m]))-- p95 latency per model, correctly aggregated across instanceshistogram_quantile(0.95,  sum by (model, le) (rate(llm_latency_seconds_bucket[5m])))-- mean input tokens per request: the silent-cost-regression detectorsum(rate(llm_tokens_total_sum{direction="in"}[5m]))  / sum(rate(llm_tokens_total_count{direction="in"}[5m]))-- week-over-week comparison for seasonal metricssum(rate(llm_requests_total[5m]))  / sum(rate(llm_requests_total[5m] offset 7d))

The histogram_quantile(0.95, sum by (model, le) (...)) form deserves attention. The sum by (..., le) adds up bucket counts across every instance before computing the quantile. If you instead computed a p95 per instance and averaged them, you would get a number with no statistical meaning — percentiles cannot be averaged. Keeping le in the by clause is the whole trick, and leaving it out is a silent, common bug.

Expensive expressions that dashboards and alerts both use should become recording rules, evaluated once on a schedule and stored as new series:

Text
groups:  - name: llm    interval: 30s    rules:      - record: llm:error_ratio_5m        expr: sum by (model) (rate(llm_requests_total{status="error"}[5m]))              / sum by (model) (rate(llm_requests_total[5m]))      - record: llm:latency_p95_5m        expr: histogram_quantile(0.95,                sum by (model, le) (rate(llm_latency_seconds_bucket[5m])))

Beyond the query-speed win, this gives you one authoritative definition of "error ratio". Without recording rules, a dashboard using a 5-minute window and an alert using a 1-minute window will disagree during every incident, and someone will spend twenty minutes discovering that the tools are both right.

Grafana: dashboards people actually use

Grafana queries other systems; it stores nothing itself. The interesting failure here is organisational, not technical: teams build a 40-panel dashboard, and during an incident nobody can find anything on it.

What works is a small hierarchy:

  • One overview dashboard, six to eight panels, answering only "is the service healthy?" — request rate, error ratio, p50/p95/p99 latency, tokens per request, cost per hour, saturation. If a panel would not change what you do in the next five minutes, it does not belong here.
  • Drill-down dashboards per subsystem, linked from the overview, where the depth lives.
  • Template variables for model, environment, and region, so one dashboard serves every deployment instead of being copy-pasted six times and diverging.
  • Annotations for deploys. A vertical line at each release turns "when did this start?" into a one-glance answer, and it is the single highest-value thing you can add to an existing dashboard.
JSON
{  "templating": {"list": [    {"name": "model", "type": "query",     "query": "label_values(llm_requests_total, model)", "includeAll": true}  ]},  "panels": [    {"title": "p95 latency", "type": "timeseries",     "targets": [{"expr": "llm:latency_p95_5m{model=~\"$model\"}"}]},    {"title": "Error ratio", "type": "timeseries",     "targets": [{"expr": "llm:error_ratio_5m{model=~\"$model\"}"}],     "fieldConfig": {"defaults": {"unit": "percentunit", "max": 0.05}}}  ]}

Log stores: index everything, or index almost nothing

The two dominant designs make opposite bets, and picking the wrong one is expensive.

Elasticsearch / OpenSearch (ELK)Loki
IndexesEvery field, into an inverted indexOnly a small set of labels; log bodies are compressed chunks
QueryFast full-text over any fieldLabel filter first, then brute-force grep over matching chunks
StorageRoughly 1x to 1.5x raw size after indexingRoughly 0.1x to 0.2x raw — no index to store
Cost driverIndex size, so CPU and SSDObject storage, which is nearly free
Bad atCost at volume; cluster operations are a real jobAd-hoc search across a wide time range with no label filter
Choose whenYou genuinely search unstructured text oftenYou mostly filter by service/level/trace ID and then read

The ELK pipeline is Filebeat (ship) → Logstash or an ingest pipeline (parse and enrich) → Elasticsearch (index) → Kibana (query). The operational lever that matters most is hot–warm–cold tiering with an index lifecycle policy: seven days on fast SSD nodes, thirty days on cheaper spinning disk with fewer replicas, then a searchable snapshot in object storage. Skipping this is how a log cluster's cost becomes linear in retention forever.

The other lever is index templates. Elasticsearch's dynamic mapping will happily create a new field for every distinct key it sees, so if you log a map keyed by user ID, you get a mapping explosion that eventually refuses writes. Set dynamic: strict or false on production indices and declare your fields.

Commercial APM: you are buying a pricing model

Datadog and New Relic both do metrics, logs, traces, and dashboards competently. The durable difference is what they charge for, because that determines which of your engineering decisions become expensive.

DatadogNew Relic
Primary unitPer host, per product (Infra, APM, Logs, Synthetics each priced separately)Per GB of telemetry ingested, plus per-seat for full-platform users
PunishesCustom-metric cardinality; running many small hostsVerbose logging; large numbers of full-access engineers
RewardsFew large hosts; disciplined taggingAggressive sampling and log trimming; small admin team
Nasty surpriseA tag added in a PR multiplies billable series (the opening story)A DEBUG flag left on in production multiplies ingest tenfold
StrengthBreadth of integrations; strong incident toolingSingle unified data model; free tier with 100 GB/month

Both publish list prices that change; treat any specific figure you read as illustrative and check the current rates. What does not change is the shape. On Datadog, cost is roughly hosts × products + cardinality, so the dangerous engineering decision is adding a label. On New Relic, cost is roughly bytes + seats, so the dangerous decision is a log level. Knowing which lever your vendor has attached to your credit card tells you which code review comment to write.

Two guardrails worth setting on day one

  1. A cardinality budget in CI. A test that instantiates your metric definitions, enumerates the declared label values, and fails the build if the product exceeds a limit. Cheap to write, and it catches the customer_id pull request before it merges rather than at the end of the billing month.
  2. A monthly spend alert on the observability bill itself. Observability is the one system where the monitoring is not monitored, and cost regressions here are discovered by finance rather than engineering.

The ML and LLM layer

Generic APM tools have no concept of a prompt, a completion, an eval score, or a distribution shift. That gap is what the ML-specific tools fill, and they split into three groups that are frequently confused.

CategoryExamplesAnswersNot for
Experiment trackingWeights & Biases, MLflow"Which hyperparameters produced the best validation loss?" Run comparison, artefact and dataset versioning, model registryLive production traffic. These are built around runs, not requests
Production ML monitoringEvidently, Arize, Fiddler"Has the input distribution moved? Is accuracy degrading? Which slice is worst?" Drift reports, performance over delayed labels, slice analysisSub-second infrastructure alerting
LLM observabilityLangSmith, Langfuse, Arize Phoenix"What was the exact prompt? Which chain step burned the tokens? Did this prompt change improve the eval score?"Host metrics and infrastructure health

This market moves fast, so check that a vendor is still actively developed before you adopt it. In the year to September 2026, for example, Neptune's hosted service shut down after OpenAI acquired the company, Helicone moved to maintenance mode after joining Mintlify, and Langfuse was acquired by ClickHouse (it stays open source and self-hostable). Data you can export in an open format, such as OpenTelemetry traces, is what makes a vendor change survivable.

The confusion worth clearing up is the first row. Weights & Biases is superb at training-time work and is repeatedly mistaken for a production monitor. The distinction is structural: an experiment tracker is organised around a run — a bounded process with a config, a set of logged scalars over steps, and artefacts. Production is organised around a request — unbounded, continuous, and needing per-slice aggregation with sub-minute freshness. You can log production aggregates into W&B, and teams do, but you are using a lab notebook as a control room.

What LLM-native tools add that APM cannot

A generic trace records that llm.call took 3.9 seconds. An LLM observability tool records the rendered prompt, the completion, the token counts, the model version, the tool calls the model requested, and — critically — attaches all of it to an eval dataset, so you can:

  • Capture real production requests that went wrong and promote them into a regression suite, one click.
  • Re-run that suite against a new prompt or model and see a per-example diff, not just an aggregate score.
  • Score outputs with an LLM judge or a rule, and track that score as a first-class metric over time.
  • Link a thumbs-down from a user back to the exact prompt, retrieved context, and completion that caused it.
Python
from langsmith import traceable@traceable(run_type="chain", name="answer_question")def answer(question: str, user_tier: str):    docs = retrieve(question, k=5)            # nested run, auto-captured    out = llm.invoke(build_prompt(question, docs))    return {        "answer": out.content,        "sources": [d.id for d in docs],        "metadata": {"user_tier": user_tier, "n_docs": len(docs)},    }

The prompts-and-completions capture is also the reason these tools need a data-governance conversation before adoption. You are, by design, sending user-submitted text to a third party. Redaction, retention limits, and a self-hosted option (Langfuse and Phoenix both offer one) are the things to check first, not the dashboard screenshots.

OpenTelemetry: the part that stops you being trapped

Every vendor above wants you to use their agent and their SDK. Do that and switching vendors means re-instrumenting every service, which is why teams stay on tools they have outgrown.

OpenTelemetry breaks that link. It is a vendor-neutral specification plus SDKs plus a Collector: your code emits OTLP, the Collector receives it, processes it, and exports it wherever you point it. Changing backend becomes a configuration change.

Text
receivers:  otlp: {protocols: {grpc: {endpoint: 0.0.0.0:4317}}}processors:  batch: {timeout: 5s, send_batch_size: 512}  attributes:                       # strip PII before it leaves your network    actions:      - {key: gen_ai.input.messages,      action: delete}      - {key: gen_ai.output.messages,     action: delete}      - {key: gen_ai.system_instructions, action: delete}      - {key: user.email,                 action: delete}  tail_sampling:    decision_wait: 30s    policies:      - {name: errors,  type: status_code, status_code: {status_codes: [ERROR]}}      - {name: slow,    type: latency,     latency: {threshold_ms: 2000}}      - {name: sample,  type: probabilistic, probabilistic: {sampling_percentage: 1}}exporters:  prometheus: {endpoint: 0.0.0.0:8889}  otlp/traces: {endpoint: tempo:4317}  otlphttp/vendor: {endpoint: https://otlp.vendor.example/v1/traces}service:  pipelines:    traces:  {receivers: [otlp], processors: [attributes, tail_sampling, batch], exporters: [otlp/traces, otlphttp/vendor]}    metrics: {receivers: [otlp], processors: [batch], exporters: [prometheus]}

Three things that configuration is doing beyond routing. It strips prompt and completion text (the attributes the OpenTelemetry GenAI conventions use for message content) and email addresses before data leaves your network, which is far safer than trusting a vendor-side redaction setting. It applies tail sampling centrally, so the policy lives in one place rather than in every service. And it exports traces to two destinations at once, which is exactly how you run a real vendor evaluation: dual-write for a month, compare, then delete one exporter line. A processor only runs if it is listed in the pipeline, so check that line: an attributes block defined above but missing from processors: redacts nothing. Order matters too: redact first, sample next, batch last.

The LLM platforms are part of this picture now. LangSmith, Langfuse and Arize Phoenix all accept OTLP traces (Langfuse over HTTP only, not gRPC), and they show spans that carry the gen_ai.* attributes as model calls with tokens and cost. So one set of OpenTelemetry instrumentation can feed your APM backend and your LLM tool at the same time.

Choosing a stack

Small team, early productML team with models in productionEnterprise / regulated
MetricsPrometheus + Grafana (or a managed Grafana Cloud free tier)Prometheus + Thanos or Mimir for long retentionManaged platform with SSO, audit logs, RBAC
LogsLoki, or the platform's built-in logsLoki or Elasticsearch with tieringElasticsearch or Splunk with retention policy per data class
TracesOTel SDK → Tempo or JaegerOTel Collector with tail sampling → TempoOTel Collector with PII stripping → vendor APM
Drift & model qualityEvidently reports on a nightly cronEvidently or a hosted ML-monitoring platform, wired to the model registryPlatform with model governance and lineage
LLM layerLangfuse or Phoenix, self-hostedLangSmith or Langfuse with eval datasets in CISelf-hosted LLM observability inside the network boundary
Typical monthly spendUnder a few hundred dollarsLow thousandsTens of thousands, plus a team that owns it

Instrument with OpenTelemetry, route through a Collector, and treat every backend as replaceable. The instrumentation is the asset you are building; the vendor is a configuration line.

How to actually make the decision

Feature matrices are close to useless here, because every product ticks every box. Four questions discriminate, and they are all about your constraints rather than the tool's capabilities.

What is your cardinality going to be in a year? Enumerate your metrics and their label values and multiply. If the answer is under a hundred thousand series, self-hosted Prometheus on one machine handles it on 26 GB of disk a month and you can stop shopping. If it is in the millions, you have either a scaling problem or — far more likely — a design problem to fix before you buy anything.

Who is on call, and how many of them are there? A three-person team cannot operate an Elasticsearch cluster and should not try; the correct choice is whatever they do not have to run. A platform team of fifteen has a different calculation, and the operational cost of self-hosting is often lower than the licence.

Can user text leave your network? If the answer is no — healthcare, finance, EU personal data — this eliminates a large fraction of the LLM observability market immediately, and it is much cheaper to discover that before integration than during a security review.

What is the exit cost? Assume you will be wrong about the vendor. If your instrumentation is OpenTelemetry, leaving costs an afternoon of Collector configuration. If it is a proprietary agent, leaving costs a quarter of engineering time, which in practice means you never leave.

The team from the opening story did not have a tooling problem. They had a metric with 800 label values and no cardinality budget in CI. The fix was a twelve-line test and moving customer_id from a metric label to a log field — after which the graph they wanted was still available, and cost nothing.