Course Content
AI Monitoring and Observability
3 sections · 7 lessons
Alerting and Automated Retraining
In one quarter — 91 days — a team's alerting channel received 4,312 messages. That is 47 a day. Reviewing them afterwards, exactly 11 corresponded to something a human should have acted on. A precision of 0.26%.
On 14 June the service began returning empty completions. The alert fired at 02:14. It was acknowledged at 09:40, when someone came into the office. Twenty-two minutes of the outage were still in progress at that point; the rest had self-healed. The alert had done its job perfectly and arrived as message number 61 that night.
This is the normal end state of alerting built by well-intentioned engineers, and it is not a discipline problem. It is an arithmetic problem. Every threshold you set is a statistical bet, and if you do not compute the false-alarm rate that bet implies, the total will be a number like 47 a day — at which point the team stops reading, and your monitoring has become strictly worse than none, because it also consumes attention.
What an alert is for
An alert exists to interrupt a human. That is expensive — it costs sleep, focus, and a slice of the team's willingness to trust the next one. So the bar is high, and there is a single test that filters most bad alerts:
If this fires at 3 a.m., is there something a human must do right now that cannot wait until morning? If not, it is not a page. It might still be a ticket, a dashboard, or nothing.
| Bad alert | Why it fails | What it should be |
|---|---|---|
| CPU above 80% | Not user-visible. High CPU on a healthy service is efficiency | A dashboard panel, or a capacity ticket |
| Any 5xx response | Fires on a single transient. Nothing to do about one | Error ratio above a threshold, sustained |
| Disk 70% full | Days of runway. Waking someone achieves nothing | Predicted to fill within 4 hours → page; otherwise ticket |
| Model accuracy dropped | Nobody can fix accuracy at 3 a.m. | Ticket for the ML team, with the drift diagnosis attached |
| Latency above 3σ of the last hour | No stated consequence; and the arithmetic below | p95 above the SLO threshold, sustained |
| Deploy completed | Not a problem at all | A deploy annotation on dashboards |
Two structural principles sit behind that table. Alert on symptoms, not causes. Users experience "answers are slow" and "answers are wrong"; they do not experience "connection pool utilisation is 94%". Symptom alerts are few, stable, and catch causes nobody predicted. Cause alerts multiply without bound and each one covers exactly one scenario. And every alert needs a runbook link — the first three things to check. An alert whose entire content is "ErrorRateHigh" hands the responder a puzzle instead of a task.
The arithmetic of a threshold
Now the part that decides whether any of this survives. Suppose you evaluate a metric once a minute. Over a week:
60 x 24 x 7 = 10,080 checks per weekAssume for the moment the metric is roughly normal and stable, and you alert whenever it exceeds the mean by k standard deviations on the upper side. The probability of one check breaching is the upper-tail area of the standard normal, 1−Φ(k). Multiply by 10,080 and you get the number of pages per week from a system where nothing whatsoever is wrong:
| Threshold | 1−Φ(k) per check | False alarms per week | In practice |
|---|---|---|---|
| 2.0σ | 0.022750 | 229.3 | 33 a day. Channel is muted within a week |
| 2.5σ | 0.0062097 | 62.6 | 9 a day. Still unusable |
| 3.0σ | 0.0013499 | 13.6 | 2 a day. The "rigorous" default, and still far too many |
| 3.5σ | 0.00023263 | 2.35 | Borderline tolerable for one metric |
| 4.0σ | 0.000031671 | 0.32 | One every 3.1 weeks — but very slow to detect real problems |
The 3σ row is the important one, because "three sigma" is what people reach for when they want to sound careful. It produces about fourteen false pages a week from a single healthy metric. And you are not watching one metric. With 40 alerting rules each set at 3σ:
40 x 13.6 = 544 false alarms per week = 78 per dayThat is the same order of noise as the 47-a-day channel from the opening, and it is produced entirely by people doing what looked like the responsible thing.
The instinct at this point is to tighten the threshold. Go to 4.5σ and you get 0.034 false alarms per week per rule, or about 1.4 a week across 40 rules. Tolerable. But you have paid for it: at 4.5σ, a real regression that shifts the mean by 2σ requires the metric to reach 2.5σ above its new mean before you hear about it, and the expected wait is 1/(1−Φ(2.5))=161 checks — nearly three hours. You have traded a noisy monitor for a blind one.
Persistence beats tightness
There is a second lever, and it is far more powerful than the threshold: require the breach to persist. Fire only when two consecutive checks both exceed the threshold.
If checks were independent, the probability that a specific adjacent pair both breach is p2. In 10,080 checks there are 10,079 adjacent pairs, so the expected number of two-in-a-row firings per week is 10,079×p2:
| Threshold | p | p2 | False alarms/week, single | False alarms/week, two consecutive |
|---|---|---|---|---|
| 2.0σ | 0.022750 | 5.176e−4 | 229.3 | 5.22 |
| 2.5σ | 0.0062097 | 3.856e−5 | 62.6 | 0.389 (one per 2.6 weeks) |
| 3.0σ | 0.0013499 | 1.822e−6 | 13.6 | 0.0184 (one per 54 weeks) |
Read the top row against the third. A loose 2σ threshold with a two-consecutive rule produces 5.2 false alarms a week. A tight 3σ threshold on a single check produces 13.6. The looser rule is 2.6 times quieter than the tighter one — because squaring a small probability is a much stronger lever than pushing further into the tail.
But quietness is only half the question. The rule that never fires is quietest of all. What does persistence cost in detection speed?
Comparing at a matched false-alarm budget
The only fair comparison holds the false-alarm rate constant and asks which rule detects faster. Fix the budget at one false page per week per rule and solve for the threshold each design needs:
single breach: 10,080 x p = 1 -> p = 9.92e-5 -> k = 3.72 sigmatwo consecutive: 10,079 x p^2 = 1 -> p = 9.96e-3 -> k = 2.33 sigmaNow suppose a real regression arrives and shifts the mean by δ standard deviations. Each check now breaches with probability 1−Φ(k−δ). For the single-breach rule the wait is geometric, so the expected number of checks to fire is 1/p. For the two-consecutive rule the expected wait until two successes in a row is (1+p)/p2.
| Real shift | Single breach at 3.72σ | Two consecutive at 2.33σ | ||
|---|---|---|---|---|
| p per check | Checks to fire | p per check | Checks to fire | |
| 2.0σ | 0.0427 | 23.4 | 0.3707 | 9.97 |
| 3.0σ | 0.2358 | 4.24 | 0.7486 | 3.12 |
| 4.0σ | 0.6103 | 1.64 | 0.9525 | 2.15 |
At identical false-alarm rates, the two-consecutive rule detects a modest 2σ regression in 10 checks instead of 23 — 2.3 times faster. For a 3σ shift it is 26% faster. Only for a huge 4σ jump does the single threshold win, by half a check, which at a one-minute cadence is 31 seconds and does not matter because a 4σ jump will trigger everything you own anyway.
A loose threshold that must persist beats a tight threshold that fires instantly, on both axes at once: fewer false alarms and faster detection of the small regressions that actually slip through.
The reason is structural. A tight single threshold has to distinguish signal from noise using one sample, so it must sit far out in the tail where real regressions rarely reach. A persistence rule gets to use the fact that noise does not repeat and problems do, which lets it sit close to the mean where regressions live.
How much persistence is too much
If two is good, is five better? No — the detection cost grows fast. The expected number of checks to obtain r consecutive successes at per-check probability p is (1−pr)/(pr(1−p)). At a 3σ threshold facing a genuine 3σ shift, p=0.5:
| Consecutive breaches required | False alarms/week | Checks to detect a 3σ shift |
|---|---|---|
| 1 | 13.6 | 2 |
| 2 | 0.018 | 6 |
| 3 | 0.000025 (one per 775 years) | 14 |
| 5 | effectively never | 62 |
Going from 1 to 2 removes 99.9% of false alarms for four extra checks. Going from 2 to 5 removes essentially nothing more — the rate was already negligible — and costs 56 additional checks, turning a six-minute detection into an hour. Two, occasionally three, is the whole range worth using. In Prometheus this is the for: clause; with a 1-minute evaluation interval, for: 2m is the two-consecutive rule.
The assumption that will bite you
All of the above assumes consecutive checks are independent. Real metrics are autocorrelated: if latency is elevated this minute it is more likely to be elevated next minute. The true joint probability lies somewhere between p2 (independent) and p (perfectly correlated), so a real system's false-alarm reduction is smaller than the table suggests.
Persistence still helps enormously, because the things it filters best — a scrape that timed out, a garbage-collection pause, a single slow request in a low-traffic window — are exactly the uncorrelated events. But do not ship a threshold derived from theory. Backtest it: replay 30 days of stored metric history through the proposed rule and count how many times it would have fired. That number is the truth; the arithmetic above is how you decide which two or three candidates to backtest.
The other assumption to check is normality. Latency is right-skewed with a hard floor at zero, so its standard deviation is a poor description of its tail and a 3σ threshold on raw latency is close to meaningless. For latency, alert on a percentile against a fixed, business-derived threshold — "p95 above 2 seconds for 5 minutes" — not on sigmas. Sigma-based rules belong on quantities that really are roughly symmetric: token counts, confidence scores, prediction-class shares, ratios.
Severity, and what each level means
| Severity | Definition | Route | Response time | Example |
|---|---|---|---|---|
| P1 / page | Users are affected now, and it will not self-heal | PagerDuty, phone | Minutes | Error ratio above 5% for 2 minutes; all completions empty |
| P2 / urgent ticket | Degraded, or will become P1 within hours | Slack channel with an owner, business hours | Hours | p95 latency above SLO for 15 minutes; error budget burning at 6x |
| P3 / ticket | Needs attention this week | Issue tracker | Days | PSI above 0.25 on an input feature; accuracy down 2 points |
| Info | Context, never an interruption | Dashboard annotation only | None | Deploy completed; retraining job finished |
The failure mode here has a name: severity inflation. Every alert author believes their alert matters, so everything becomes P1, and P1 stops meaning anything. A useful forcing function is a budget — no more than, say, twelve P1 rules for the whole service — which turns "should this page?" into a comparison against the existing twelve rather than a solo judgement.
Error budgets and burn-rate alerting
The most robust alerting scheme available ties the threshold to a business commitment rather than to a distribution. Start from an SLO: 99.9% of requests succeed over 30 days. That grants an error budget:
30 days = 43,200 minutesbudget = 43,200 x 0.001 = 43.2 minutes of failure per monthDefine the burn rate as the observed error ratio divided by the budgeted one. Burning at 1x consumes the budget in exactly 30 days. Burning at 14.4x consumes it in 30/14.4 = 2.08 days. The alert then encodes both severity and urgency in one number, and the standard multi-window scheme is:
| Burn rate | Long window | Short window | Budget consumed | Action |
|---|---|---|---|---|
| 14.4x | 1 hour | 5 min | 2% in one hour | Page |
| 6x | 6 hours | 30 min | 5% in six hours | Page |
| 3x | 1 day | 2 hours | 10% in a day | Ticket |
| 1x | 3 days | 6 hours | 10% in three days | Ticket |
Check the top row: one hour is 1/720 of a 30-day month, so at 14.4x you consume 14.4/720 = 2.0% of the budget in that hour. The short window is the trick that makes this responsive — the alert requires both windows to be burning, so it resolves quickly once the incident ends instead of staying lit for an hour after recovery.
groups: - name: slo rules: - alert: ErrorBudgetBurnFast expr: | (llm:error_ratio_1h > 14.4 * 0.001) and (llm:error_ratio_5m > 14.4 * 0.001) labels: {severity: page} annotations: summary: "Burning error budget 14.4x - 2% of the month in one hour" runbook: "https://wiki.internal/runbooks/llm-error-budget" - alert: LatencySLOBreach expr: llm:latency_p95_5m > 2.0 for: 5m # five consecutive evaluations labels: {severity: page} annotations: summary: "p95 latency {{ $value | humanizeDuration }} exceeds 2s SLO" - alert: InputDriftHigh expr: model_feature_psi > 0.25 for: 2h # drift is slow; do not page on a blip labels: {severity: ticket} annotations: summary: "PSI {{ $value }} on {{ $labels.feature }} vs training baseline"Routing: turning 40 alerts into one page
A single bad deploy will trip a dozen rules at once. Alertmanager's job is to make that one interruption rather than twelve, and three features do the work.
- Grouping. Collapse alerts sharing labels — service, cluster, alertname — into one notification, held briefly so that related alerts arrive together rather than as a burst.
- Inhibition. Suppress downstream alerts when an upstream one is firing. If
ServiceDownis active, do not also page aboutHighLatencyandErrorRateHighfor the same service — they are consequences, and the responder already knows. - Silences. Time-bounded mutes with an author and a reason, applied during known maintenance. Bounded is the key word; a silence with no expiry is how an alert disappears permanently and nobody notices for a year.
route: group_by: [alertname, service] group_wait: 30s # let related alerts arrive before notifying group_interval: 5m repeat_interval: 4h receiver: slack-default routes: - matchers: [severity="page"] receiver: pagerduty continue: true # also post to Slack for visibility - matchers: [severity="ticket"] receiver: jirainhibit_rules: - source_matchers: [alertname="ServiceDown"] target_matchers: [severity=~"page|ticket"] equal: [service]Automated retraining, and the guardrails that make it safe
The tempting wiring is "drift alert fires → retraining pipeline starts → new model deploys". Every part of that chain has a specific way of going wrong.
| Trigger | Sensible | Danger |
|---|---|---|
| Scheduled (weekly, monthly) | Predictable; easy to reason about; catches slow drift | Retrains when nothing changed, burning compute and review time |
| Performance-based (accuracy below threshold on labels) | The most defensible trigger — it responds to the thing you care about | Needs labels, which arrive late; the damage is done by the time it fires |
| Drift-based (PSI above 0.25) | Fast; fires before labels exist | The most common cause of large PSI is an upstream data bug |
| Volume-based (50,000 new labelled examples) | Ties retraining to genuinely new information | New data may be biased towards whatever the current model surfaced |
That third row is the one that turns an outage into a permanent regression. Recall the pattern: a partner starts sending distances in miles instead of kilometres, PSI spikes, the automated pipeline retrains on the corrupted feed, and the corruption is now in the weights. Fixing the pipeline no longer fixes the model. Drift opens a ticket. A human decides whether the answer is a retrain, a pipeline fix, or nothing.
Do the cost arithmetic too. Suppose a retrain costs 400 dollars of compute plus four engineer-hours of validation. A drift rule that produces one false trigger a week costs:
52 x 400 dollars = 20,800 per year in compute52 x 4 engineer-hours = 208 hours per year of reviewFive weeks of an engineer's year spent validating retrains that were never needed. Alert thresholds have a direct line to a budget.
The gate a challenger must pass
When a retrain does happen, "the new model is better" needs a number, because accuracy measured on a finite holdout is itself noisy. With a holdout of 2,000 examples and a champion accuracy of 0.91, the standard error of a single accuracy estimate is:
SE = sqrt(0.91 x 0.09 / 2000) = sqrt(4.095e-5) = 0.0064 = 0.64 percentage pointsSo a challenger scoring 91.5% is +0.5 pp — comfortably inside the noise of a single measurement, and completely meaningless as evidence. A defensible gate requires the improvement to clear roughly two standard errors, so about +1.5 pp on this holdout, or better still uses a paired comparison (McNemar's test on the examples where the two models disagree), which is far more sensitive because it removes the variance shared by both models.
1def promotion_gate(champion, challenger, holdout, slices):2 checks = {3 # headline metric must beat the incumbent by more than measurement noise4 "beats_champion": challenger.accuracy - champion.accuracy > 0.015,5 # no slice may regress, however good the average looks6 "no_slice_regression": all(7 challenger.accuracy_on(s) >= champion.accuracy_on(s) - 0.018 for s in slices),9 # the known-failure regression suite must not break10 "regression_suite": challenger.score_on(REGRESSION_SET) >= 0.95,11 # trained on data that passed validation12 "training_data_clean": challenger.training_data.validation_passed,13 # cost and latency did not silently blow out14 "latency_ok": challenger.p95_latency_ms <= champion.p95_latency_ms * 1.10,15 }16 return all(checks.values()), checksThe no_slice_regression check earns its place repeatedly. A model can gain 2 points overall while losing 9 points on a minority segment, because the aggregate is dominated by the majority. Aggregate metrics hide exactly the failures that generate complaints.
Beyond the gate, the deployment itself should be staged: run the challenger in shadow first (it sees live traffic, its outputs are logged and never served), then canary to 5% of traffic with automatic rollback on any guardrail breach, then ramp. And keep the previous model artefact loadable, with a rollback that is one command and takes under five minutes. Automated promotion without automated rollback is not automation; it is a faster way to break production.
What this means when you build one
Three habits separate an alerting system that survives a year from one that becomes a muted channel.
Compute the false-alarm rate before you merge the rule. Number of checks per week times the tail probability, times the number of rules. If the total across your whole alert set is more than about five per week, you are building the 47-a-day channel and you will get there in about a quarter. This calculation takes two minutes and almost nobody does it.
Reach for persistence before you reach for a tighter threshold. At a matched false-alarm budget, a 2.33σ threshold requiring two consecutive breaches detects a 2σ regression in 10 checks where a 3.72σ single-breach rule takes 23. Then backtest both against 30 days of real history, because your metric is autocorrelated and the theoretical numbers are optimistic.
Review the alerts, not just the incidents. Once a month, list every alert that fired and mark each one: acted on, ignored, or auto-resolved. Anything in the "ignored" column for two consecutive months gets deleted or downgraded — no exceptions, no "but it might be useful someday". An alert nobody acts on is not free; it is spending the attention that the next real page will need.
The team from the opening did exactly this. They deleted 31 of their 40 rules outright, converted 6 into dashboard panels, rewrote the remaining 3 as burn-rate alerts with two-window conditions, and added a drift rule routed to a ticket queue rather than a pager. The following quarter their channel received 38 messages. Nine of them were real. The 02:14 alert would now be the only thing in the channel.