Course Content
AI Monitoring and Observability
3 sections · 7 lessons
Monitoring Tools and Platforms Overview
A twelve-person team ships an LLM product. Their month-three observability bill is 2,800 dollars — twenty hosts of infrastructure and APM monitoring, plus logs. Reasonable. Month four the bill is 9,100 dollars. Traffic grew 6%.
The cause is one line in a pull request. An engineer wanted per-customer latency graphs, so they added a customer_id tag to four metrics. Those metrics already carried a model label (5 values) and an endpoint label (8 values), so they represented 4 × 5 × 8 = 160 time series. The company has 800 customers:
160 series x 800 customers = 128,000 custom metricsincluded with 20 hosts (100 each) = 2,000billable = 126,000at 0.05/metric/month = 6,300/monthSix thousand three hundred dollars a month for a graph that was viewed twice. The same information was already sitting in their logs, sliceable for free.
Choosing observability tools is not really a feature comparison. Every product in this space can draw a line chart. What differs is the cost model, the data model, and what happens when you want to leave — and those three things determine whether the stack you pick survives contact with growth.
Four jobs, and why no single tool does all of them
People say "we use Datadog" or "we use Prometheus" as though it were one decision. It is four, and the four have genuinely different storage requirements.
| Job | Data shape | Query pattern | Scales with | Typical tools |
|---|---|---|---|---|
| Metrics store | Numeric time series, fixed labels | Aggregate over time, group by label | Active series (cardinality) | Prometheus, Mimir, VictoriaMetrics, Datadog |
| Log store | Semi-structured events, arbitrary fields | Filter, full-text search, count | Bytes ingested and indexed | Elasticsearch/OpenSearch, Loki, Datadog Logs, Splunk |
| Trace store | Span trees, high cardinality by nature | Fetch one trace; aggregate over span attributes | Spans stored (after sampling) | Jaeger, Tempo, Datadog APM, New Relic |
| Model / LLM layer | Prompts, completions, scores, datasets, drift stats | Compare runs, inspect examples, score evals | Requests logged, eval runs | LangSmith, Langfuse, Arize Phoenix, Evidently |
A metrics database that also indexes arbitrary text is a search engine wearing a costume, and it will be bad at one of the two. The reason "one pane of glass" products exist is that switching context during an incident is genuinely expensive — but internally they are still running four different storage engines, and you pay for each.
Prometheus: the pull model and what it buys you
Prometheus is the default open-source metrics store, and its most distinctive design choice is that it pulls. Your application exposes a plain-text endpoint at /metrics; Prometheus scrapes it on a schedule.
# HELP llm_requests_total Total LLM requests# TYPE llm_requests_total counterllm_requests_total{model="assistant-v4",endpoint="/chat",status="ok"} 148203llm_requests_total{model="assistant-v4",endpoint="/chat",status="error"} 71# HELP llm_latency_seconds End-to-end latency# TYPE llm_latency_seconds histogramllm_latency_seconds_bucket{model="assistant-v4",le="0.5"} 91442llm_latency_seconds_bucket{model="assistant-v4",le="1.0"} 132890llm_latency_seconds_bucket{model="assistant-v4",le="+Inf"} 148274llm_latency_seconds_sum{model="assistant-v4"} 96331.4llm_latency_seconds_count{model="assistant-v4"} 148274Pull sounds like an implementation detail and is not. Four consequences:
- Scrape failure is itself a signal. If a target stops responding, Prometheus records
up = 0. With push, a silent instance is indistinguishable from a healthy quiet one. - Back-pressure is automatic. An overloaded Prometheus scrapes more slowly. It does not get flooded by applications pushing harder.
- You can run the endpoint by hand.
curl localhost:8000/metricsduring development shows exactly what will be collected, which makes instrumentation bugs obvious in seconds. - Short-lived jobs do not fit. A batch job that runs for 40 seconds may never be scraped. That is what the Pushgateway is for — and it is a genuine exception, not a general-purpose push endpoint. Metrics pushed there persist until deleted, so using it for ordinary services gives you stale values forever after a deploy.
The architecture is small: a scrape loop feeding a local time-series database, a rule evaluator computing recording and alerting rules on a schedule, an HTTP query API, and Alertmanager as a separate process handling grouping, silencing, and routing.
Retention arithmetic, so you can size it
Prometheus compresses aggressively — roughly 1.5 bytes per sample after delta-of-delta and XOR encoding. With 100,000 active series and a 15-second scrape interval:
samples/sec = 100,000 / 15 = 6,667samples/day = 6,667 x 86,400 = 576,000,000bytes/day = 576,000,000 x 1.5 = 864 MB/day30-day disk = 864 MB x 30 = 25.9 GBTwenty-six gigabytes for a month of a hundred thousand series. This is why a single Prometheus on a modest box goes a very long way, and why the framing "we need a managed metrics platform" often means "we have a cardinality problem" rather than "we have a scale problem".
Watch what the earlier tagging mistake does to the same calculation at 12 million series:
samples/sec = 12,000,000 / 15 = 800,000samples/day = 800,000 x 86,400 = 69,120,000,000bytes/day = 69.12e9 x 1.5 = 103.7 GB/dayRAM for the active series head block (~3.5 KB each) = ~42 GBThe disk is survivable. The 42 GB of resident memory is not, and this is how Prometheus servers die: not gradually, but by out-of-memory kill during a deploy that adds one label.
Prometheus scales beautifully with time and traffic and catastrophically with cardinality. Every capacity conversation about a metrics store is really a conversation about label values.
PromQL, and recording rules
1-- error ratio over 5 minutes2sum(rate(llm_requests_total{status="error"}[5m]))3 / sum(rate(llm_requests_total[5m]))45-- p95 latency per model, correctly aggregated across instances6histogram_quantile(0.95,7 sum by (model, le) (rate(llm_latency_seconds_bucket[5m])))89-- mean input tokens per request: the silent-cost-regression detector10sum(rate(llm_tokens_total_sum{direction="in"}[5m]))11 / sum(rate(llm_tokens_total_count{direction="in"}[5m]))1213-- week-over-week comparison for seasonal metrics14sum(rate(llm_requests_total[5m]))15 / sum(rate(llm_requests_total[5m] offset 7d))The histogram_quantile(0.95, sum by (model, le) (...)) form deserves attention. The sum by (..., le) adds up bucket counts across every instance before computing the quantile. If you instead computed a p95 per instance and averaged them, you would get a number with no statistical meaning — percentiles cannot be averaged. Keeping le in the by clause is the whole trick, and leaving it out is a silent, common bug.
Expensive expressions that dashboards and alerts both use should become recording rules, evaluated once on a schedule and stored as new series:
groups: - name: llm interval: 30s rules: - record: llm:error_ratio_5m expr: sum by (model) (rate(llm_requests_total{status="error"}[5m])) / sum by (model) (rate(llm_requests_total[5m])) - record: llm:latency_p95_5m expr: histogram_quantile(0.95, sum by (model, le) (rate(llm_latency_seconds_bucket[5m])))Beyond the query-speed win, this gives you one authoritative definition of "error ratio". Without recording rules, a dashboard using a 5-minute window and an alert using a 1-minute window will disagree during every incident, and someone will spend twenty minutes discovering that the tools are both right.
Grafana: dashboards people actually use
Grafana queries other systems; it stores nothing itself. The interesting failure here is organisational, not technical: teams build a 40-panel dashboard, and during an incident nobody can find anything on it.
What works is a small hierarchy:
- One overview dashboard, six to eight panels, answering only "is the service healthy?" — request rate, error ratio, p50/p95/p99 latency, tokens per request, cost per hour, saturation. If a panel would not change what you do in the next five minutes, it does not belong here.
- Drill-down dashboards per subsystem, linked from the overview, where the depth lives.
- Template variables for model, environment, and region, so one dashboard serves every deployment instead of being copy-pasted six times and diverging.
- Annotations for deploys. A vertical line at each release turns "when did this start?" into a one-glance answer, and it is the single highest-value thing you can add to an existing dashboard.
1{2 "templating": {"list": [3 {"name": "model", "type": "query",4 "query": "label_values(llm_requests_total, model)", "includeAll": true}5 ]},6 "panels": [7 {"title": "p95 latency", "type": "timeseries",8 "targets": [{"expr": "llm:latency_p95_5m{model=~\"$model\"}"}]},9 {"title": "Error ratio", "type": "timeseries",10 "targets": [{"expr": "llm:error_ratio_5m{model=~\"$model\"}"}],11 "fieldConfig": {"defaults": {"unit": "percentunit", "max": 0.05}}}12 ]13}Log stores: index everything, or index almost nothing
The two dominant designs make opposite bets, and picking the wrong one is expensive.
| Elasticsearch / OpenSearch (ELK) | Loki | |
|---|---|---|
| Indexes | Every field, into an inverted index | Only a small set of labels; log bodies are compressed chunks |
| Query | Fast full-text over any field | Label filter first, then brute-force grep over matching chunks |
| Storage | Roughly 1x to 1.5x raw size after indexing | Roughly 0.1x to 0.2x raw — no index to store |
| Cost driver | Index size, so CPU and SSD | Object storage, which is nearly free |
| Bad at | Cost at volume; cluster operations are a real job | Ad-hoc search across a wide time range with no label filter |
| Choose when | You genuinely search unstructured text often | You mostly filter by service/level/trace ID and then read |
The ELK pipeline is Filebeat (ship) → Logstash or an ingest pipeline (parse and enrich) → Elasticsearch (index) → Kibana (query). The operational lever that matters most is hot–warm–cold tiering with an index lifecycle policy: seven days on fast SSD nodes, thirty days on cheaper spinning disk with fewer replicas, then a searchable snapshot in object storage. Skipping this is how a log cluster's cost becomes linear in retention forever.
The other lever is index templates. Elasticsearch's dynamic mapping will happily create a new field for every distinct key it sees, so if you log a map keyed by user ID, you get a mapping explosion that eventually refuses writes. Set dynamic: strict or false on production indices and declare your fields.
Commercial APM: you are buying a pricing model
Datadog and New Relic both do metrics, logs, traces, and dashboards competently. The durable difference is what they charge for, because that determines which of your engineering decisions become expensive.
| Datadog | New Relic | |
|---|---|---|
| Primary unit | Per host, per product (Infra, APM, Logs, Synthetics each priced separately) | Per GB of telemetry ingested, plus per-seat for full-platform users |
| Punishes | Custom-metric cardinality; running many small hosts | Verbose logging; large numbers of full-access engineers |
| Rewards | Few large hosts; disciplined tagging | Aggressive sampling and log trimming; small admin team |
| Nasty surprise | A tag added in a PR multiplies billable series (the opening story) | A DEBUG flag left on in production multiplies ingest tenfold |
| Strength | Breadth of integrations; strong incident tooling | Single unified data model; free tier with 100 GB/month |
Both publish list prices that change; treat any specific figure you read as illustrative and check the current rates. What does not change is the shape. On Datadog, cost is roughly hosts × products + cardinality, so the dangerous engineering decision is adding a label. On New Relic, cost is roughly bytes + seats, so the dangerous decision is a log level. Knowing which lever your vendor has attached to your credit card tells you which code review comment to write.
Two guardrails worth setting on day one
- A cardinality budget in CI. A test that instantiates your metric definitions, enumerates the declared label values, and fails the build if the product exceeds a limit. Cheap to write, and it catches the
customer_idpull request before it merges rather than at the end of the billing month. - A monthly spend alert on the observability bill itself. Observability is the one system where the monitoring is not monitored, and cost regressions here are discovered by finance rather than engineering.
The ML and LLM layer
Generic APM tools have no concept of a prompt, a completion, an eval score, or a distribution shift. That gap is what the ML-specific tools fill, and they split into three groups that are frequently confused.
| Category | Examples | Answers | Not for |
|---|---|---|---|
| Experiment tracking | Weights & Biases, MLflow | "Which hyperparameters produced the best validation loss?" Run comparison, artefact and dataset versioning, model registry | Live production traffic. These are built around runs, not requests |
| Production ML monitoring | Evidently, Arize, Fiddler | "Has the input distribution moved? Is accuracy degrading? Which slice is worst?" Drift reports, performance over delayed labels, slice analysis | Sub-second infrastructure alerting |
| LLM observability | LangSmith, Langfuse, Arize Phoenix | "What was the exact prompt? Which chain step burned the tokens? Did this prompt change improve the eval score?" | Host metrics and infrastructure health |
This market moves fast, so check that a vendor is still actively developed before you adopt it. In the year to September 2026, for example, Neptune's hosted service shut down after OpenAI acquired the company, Helicone moved to maintenance mode after joining Mintlify, and Langfuse was acquired by ClickHouse (it stays open source and self-hostable). Data you can export in an open format, such as OpenTelemetry traces, is what makes a vendor change survivable.
The confusion worth clearing up is the first row. Weights & Biases is superb at training-time work and is repeatedly mistaken for a production monitor. The distinction is structural: an experiment tracker is organised around a run — a bounded process with a config, a set of logged scalars over steps, and artefacts. Production is organised around a request — unbounded, continuous, and needing per-slice aggregation with sub-minute freshness. You can log production aggregates into W&B, and teams do, but you are using a lab notebook as a control room.
What LLM-native tools add that APM cannot
A generic trace records that llm.call took 3.9 seconds. An LLM observability tool records the rendered prompt, the completion, the token counts, the model version, the tool calls the model requested, and — critically — attaches all of it to an eval dataset, so you can:
- Capture real production requests that went wrong and promote them into a regression suite, one click.
- Re-run that suite against a new prompt or model and see a per-example diff, not just an aggregate score.
- Score outputs with an LLM judge or a rule, and track that score as a first-class metric over time.
- Link a thumbs-down from a user back to the exact prompt, retrieved context, and completion that caused it.
1from langsmith import traceable23@traceable(run_type="chain", name="answer_question")4def answer(question: str, user_tier: str):5 docs = retrieve(question, k=5) # nested run, auto-captured6 out = llm.invoke(build_prompt(question, docs))7 return {8 "answer": out.content,9 "sources": [d.id for d in docs],10 "metadata": {"user_tier": user_tier, "n_docs": len(docs)},11 }The prompts-and-completions capture is also the reason these tools need a data-governance conversation before adoption. You are, by design, sending user-submitted text to a third party. Redaction, retention limits, and a self-hosted option (Langfuse and Phoenix both offer one) are the things to check first, not the dashboard screenshots.
OpenTelemetry: the part that stops you being trapped
Every vendor above wants you to use their agent and their SDK. Do that and switching vendors means re-instrumenting every service, which is why teams stay on tools they have outgrown.
OpenTelemetry breaks that link. It is a vendor-neutral specification plus SDKs plus a Collector: your code emits OTLP, the Collector receives it, processes it, and exports it wherever you point it. Changing backend becomes a configuration change.
receivers: otlp: {protocols: {grpc: {endpoint: 0.0.0.0:4317}}}processors: batch: {timeout: 5s, send_batch_size: 512} attributes: # strip PII before it leaves your network actions: - {key: gen_ai.input.messages, action: delete} - {key: gen_ai.output.messages, action: delete} - {key: gen_ai.system_instructions, action: delete} - {key: user.email, action: delete} tail_sampling: decision_wait: 30s policies: - {name: errors, type: status_code, status_code: {status_codes: [ERROR]}} - {name: slow, type: latency, latency: {threshold_ms: 2000}} - {name: sample, type: probabilistic, probabilistic: {sampling_percentage: 1}}exporters: prometheus: {endpoint: 0.0.0.0:8889} otlp/traces: {endpoint: tempo:4317} otlphttp/vendor: {endpoint: https://otlp.vendor.example/v1/traces}service: pipelines: traces: {receivers: [otlp], processors: [attributes, tail_sampling, batch], exporters: [otlp/traces, otlphttp/vendor]} metrics: {receivers: [otlp], processors: [batch], exporters: [prometheus]}Three things that configuration is doing beyond routing. It strips prompt and completion text (the attributes the OpenTelemetry GenAI conventions use for message content) and email addresses before data leaves your network, which is far safer than trusting a vendor-side redaction setting. It applies tail sampling centrally, so the policy lives in one place rather than in every service. And it exports traces to two destinations at once, which is exactly how you run a real vendor evaluation: dual-write for a month, compare, then delete one exporter line. A processor only runs if it is listed in the pipeline, so check that line: an attributes block defined above but missing from processors: redacts nothing. Order matters too: redact first, sample next, batch last.
The LLM platforms are part of this picture now. LangSmith, Langfuse and Arize Phoenix all accept OTLP traces (Langfuse over HTTP only, not gRPC), and they show spans that carry the gen_ai.* attributes as model calls with tokens and cost. So one set of OpenTelemetry instrumentation can feed your APM backend and your LLM tool at the same time.
Choosing a stack
| Small team, early product | ML team with models in production | Enterprise / regulated | |
|---|---|---|---|
| Metrics | Prometheus + Grafana (or a managed Grafana Cloud free tier) | Prometheus + Thanos or Mimir for long retention | Managed platform with SSO, audit logs, RBAC |
| Logs | Loki, or the platform's built-in logs | Loki or Elasticsearch with tiering | Elasticsearch or Splunk with retention policy per data class |
| Traces | OTel SDK → Tempo or Jaeger | OTel Collector with tail sampling → Tempo | OTel Collector with PII stripping → vendor APM |
| Drift & model quality | Evidently reports on a nightly cron | Evidently or a hosted ML-monitoring platform, wired to the model registry | Platform with model governance and lineage |
| LLM layer | Langfuse or Phoenix, self-hosted | LangSmith or Langfuse with eval datasets in CI | Self-hosted LLM observability inside the network boundary |
| Typical monthly spend | Under a few hundred dollars | Low thousands | Tens of thousands, plus a team that owns it |
Instrument with OpenTelemetry, route through a Collector, and treat every backend as replaceable. The instrumentation is the asset you are building; the vendor is a configuration line.
How to actually make the decision
Feature matrices are close to useless here, because every product ticks every box. Four questions discriminate, and they are all about your constraints rather than the tool's capabilities.
What is your cardinality going to be in a year? Enumerate your metrics and their label values and multiply. If the answer is under a hundred thousand series, self-hosted Prometheus on one machine handles it on 26 GB of disk a month and you can stop shopping. If it is in the millions, you have either a scaling problem or — far more likely — a design problem to fix before you buy anything.
Who is on call, and how many of them are there? A three-person team cannot operate an Elasticsearch cluster and should not try; the correct choice is whatever they do not have to run. A platform team of fifteen has a different calculation, and the operational cost of self-hosting is often lower than the licence.
Can user text leave your network? If the answer is no — healthcare, finance, EU personal data — this eliminates a large fraction of the LLM observability market immediately, and it is much cheaper to discover that before integration than during a security review.
What is the exit cost? Assume you will be wrong about the vendor. If your instrumentation is OpenTelemetry, leaving costs an afternoon of Collector configuration. If it is a proprietary agent, leaving costs a quarter of engineering time, which in practice means you never leave.
The team from the opening story did not have a tooling problem. They had a metric with 800 label values and no cardinality budget in CI. The fix was a twelve-line test and moving customer_id from a metric label to a log field — after which the graph they wanted was still available, and cost nothing.