AI Monitoring and Observability

Alerting and Automated Retraining


In one quarter — 91 days — a team's alerting channel received 4,312 messages. That is 47 a day. Reviewing them afterwards, exactly 11 corresponded to something a human should have acted on. A precision of 0.26%.

On 14 June the service began returning empty completions. The alert fired at 02:14. It was acknowledged at 09:40, when someone came into the office. Twenty-two minutes of the outage were still in progress at that point; the rest had self-healed. The alert had done its job perfectly and arrived as message number 61 that night.

This is the normal end state of alerting built by well-intentioned engineers, and it is not a discipline problem. It is an arithmetic problem. Every threshold you set is a statistical bet, and if you do not compute the false-alarm rate that bet implies, the total will be a number like 47 a day — at which point the team stops reading, and your monitoring has become strictly worse than none, because it also consumes attention.

Persistence beats tightness at a matched budgetabove 800 ms4,31211above 800 ms64011above 700 ms9611above 650 ms2410p95 thresholdPages per quarterReal incidentsFire on one sample2 of 3 samplesSustained 5 minutesSustained 15 minutes4,312 messages in 91 days, 11 of them actionable — a precision of 0.26 percent.
Requiring the condition to hold, rather than merely to happen, keeps every true detection and deletes the noise.

What an alert is for

An alert exists to interrupt a human. That is expensive — it costs sleep, focus, and a slice of the team's willingness to trust the next one. So the bar is high, and there is a single test that filters most bad alerts:

If this fires at 3 a.m., is there something a human must do right now that cannot wait until morning? If not, it is not a page. It might still be a ticket, a dashboard, or nothing.

Bad alertWhy it failsWhat it should be
CPU above 80%Not user-visible. High CPU on a healthy service is efficiencyA dashboard panel, or a capacity ticket
Any 5xx responseFires on a single transient. Nothing to do about oneError ratio above a threshold, sustained
Disk 70% fullDays of runway. Waking someone achieves nothingPredicted to fill within 4 hours → page; otherwise ticket
Model accuracy droppedNobody can fix accuracy at 3 a.m.Ticket for the ML team, with the drift diagnosis attached
Latency above 3σ of the last hourNo stated consequence; and the arithmetic belowp95 above the SLO threshold, sustained
Deploy completedNot a problem at allA deploy annotation on dashboards

Two structural principles sit behind that table. Alert on symptoms, not causes. Users experience "answers are slow" and "answers are wrong"; they do not experience "connection pool utilisation is 94%". Symptom alerts are few, stable, and catch causes nobody predicted. Cause alerts multiply without bound and each one covers exactly one scenario. And every alert needs a runbook link — the first three things to check. An alert whose entire content is "ErrorRateHigh" hands the responder a puzzle instead of a task.

The arithmetic of a threshold

Now the part that decides whether any of this survives. Suppose you evaluate a metric once a minute. Over a week:

Text
60 x 24 x 7 = 10,080 checks per week

Assume for the moment the metric is roughly normal and stable, and you alert whenever it exceeds the mean by kk standard deviations on the upper side. The probability of one check breaching is the upper-tail area of the standard normal, 1−Φ(k)1 - \Phi(k). Multiply by 10,080 and you get the number of pages per week from a system where nothing whatsoever is wrong:

Threshold1−Φ(k)1-\Phi(k) per checkFalse alarms per weekIn practice
2.0σ0.022750229.333 a day. Channel is muted within a week
2.5σ0.006209762.69 a day. Still unusable
3.0σ0.001349913.62 a day. The "rigorous" default, and still far too many
3.5σ0.000232632.35Borderline tolerable for one metric
4.0σ0.0000316710.32One every 3.1 weeks — but very slow to detect real problems

The 3σ row is the important one, because "three sigma" is what people reach for when they want to sound careful. It produces about fourteen false pages a week from a single healthy metric. And you are not watching one metric. With 40 alerting rules each set at 3σ:

Text
40 x 13.6 = 544 false alarms per week = 78 per day

That is the same order of noise as the 47-a-day channel from the opening, and it is produced entirely by people doing what looked like the responsible thing.

The instinct at this point is to tighten the threshold. Go to 4.5σ and you get 0.034 false alarms per week per rule, or about 1.4 a week across 40 rules. Tolerable. But you have paid for it: at 4.5σ, a real regression that shifts the mean by 2σ requires the metric to reach 2.5σ above its new mean before you hear about it, and the expected wait is 1/(1−Φ(2.5))=1611/(1-\Phi(2.5)) = 161 checks — nearly three hours. You have traded a noisy monitor for a blind one.

Persistence beats tightness

There is a second lever, and it is far more powerful than the threshold: require the breach to persist. Fire only when two consecutive checks both exceed the threshold.

If checks were independent, the probability that a specific adjacent pair both breach is p2p^2. In 10,080 checks there are 10,079 adjacent pairs, so the expected number of two-in-a-row firings per week is 10,079×p210{,}079 \times p^2:

Thresholdppp2p^2False alarms/week, singleFalse alarms/week, two consecutive
2.0σ0.0227505.176e−4229.35.22
2.5σ0.00620973.856e−562.60.389 (one per 2.6 weeks)
3.0σ0.00134991.822e−613.60.0184 (one per 54 weeks)

Read the top row against the third. A loose 2σ threshold with a two-consecutive rule produces 5.2 false alarms a week. A tight 3σ threshold on a single check produces 13.6. The looser rule is 2.6 times quieter than the tighter one — because squaring a small probability is a much stronger lever than pushing further into the tail.

But quietness is only half the question. The rule that never fires is quietest of all. What does persistence cost in detection speed?

Comparing at a matched false-alarm budget

The only fair comparison holds the false-alarm rate constant and asks which rule detects faster. Fix the budget at one false page per week per rule and solve for the threshold each design needs:

Text
single breach:      10,080 x p = 1   ->  p = 9.92e-5   ->  k = 3.72 sigmatwo consecutive:    10,079 x p^2 = 1 ->  p = 9.96e-3   ->  k = 2.33 sigma

Now suppose a real regression arrives and shifts the mean by δ\delta standard deviations. Each check now breaches with probability 1−Φ(k−δ)1-\Phi(k-\delta). For the single-breach rule the wait is geometric, so the expected number of checks to fire is 1/p1/p. For the two-consecutive rule the expected wait until two successes in a row is (1+p)/p2(1+p)/p^2.

Real shiftSingle breach at 3.72σTwo consecutive at 2.33σ
pp per checkChecks to firepp per checkChecks to fire
2.0σ0.042723.40.37079.97
3.0σ0.23584.240.74863.12
4.0σ0.61031.640.95252.15

At identical false-alarm rates, the two-consecutive rule detects a modest 2σ regression in 10 checks instead of 23 — 2.3 times faster. For a 3σ shift it is 26% faster. Only for a huge 4σ jump does the single threshold win, by half a check, which at a one-minute cadence is 31 seconds and does not matter because a 4σ jump will trigger everything you own anyway.

A loose threshold that must persist beats a tight threshold that fires instantly, on both axes at once: fewer false alarms and faster detection of the small regressions that actually slip through.

The reason is structural. A tight single threshold has to distinguish signal from noise using one sample, so it must sit far out in the tail where real regressions rarely reach. A persistence rule gets to use the fact that noise does not repeat and problems do, which lets it sit close to the mean where regressions live.

How much persistence is too much

If two is good, is five better? No — the detection cost grows fast. The expected number of checks to obtain rr consecutive successes at per-check probability pp is (1−pr)/(pr(1−p))(1-p^r)/(p^r(1-p)). At a 3σ threshold facing a genuine 3σ shift, p=0.5p = 0.5:

Consecutive breaches requiredFalse alarms/weekChecks to detect a 3σ shift
113.62
20.0186
30.000025 (one per 775 years)14
5effectively never62

Going from 1 to 2 removes 99.9% of false alarms for four extra checks. Going from 2 to 5 removes essentially nothing more — the rate was already negligible — and costs 56 additional checks, turning a six-minute detection into an hour. Two, occasionally three, is the whole range worth using. In Prometheus this is the for: clause; with a 1-minute evaluation interval, for: 2m is the two-consecutive rule.

The assumption that will bite you

All of the above assumes consecutive checks are independent. Real metrics are autocorrelated: if latency is elevated this minute it is more likely to be elevated next minute. The true joint probability lies somewhere between p2p^2 (independent) and pp (perfectly correlated), so a real system's false-alarm reduction is smaller than the table suggests.

Persistence still helps enormously, because the things it filters best — a scrape that timed out, a garbage-collection pause, a single slow request in a low-traffic window — are exactly the uncorrelated events. But do not ship a threshold derived from theory. Backtest it: replay 30 days of stored metric history through the proposed rule and count how many times it would have fired. That number is the truth; the arithmetic above is how you decide which two or three candidates to backtest.

The other assumption to check is normality. Latency is right-skewed with a hard floor at zero, so its standard deviation is a poor description of its tail and a 3σ threshold on raw latency is close to meaningless. For latency, alert on a percentile against a fixed, business-derived threshold — "p95 above 2 seconds for 5 minutes" — not on sigmas. Sigma-based rules belong on quantities that really are roughly symmetric: token counts, confidence scores, prediction-class shares, ratios.

Severity, and what each level means

SeverityDefinitionRouteResponse timeExample
P1 / pageUsers are affected now, and it will not self-healPagerDuty, phoneMinutesError ratio above 5% for 2 minutes; all completions empty
P2 / urgent ticketDegraded, or will become P1 within hoursSlack channel with an owner, business hoursHoursp95 latency above SLO for 15 minutes; error budget burning at 6x
P3 / ticketNeeds attention this weekIssue trackerDaysPSI above 0.25 on an input feature; accuracy down 2 points
InfoContext, never an interruptionDashboard annotation onlyNoneDeploy completed; retraining job finished

The failure mode here has a name: severity inflation. Every alert author believes their alert matters, so everything becomes P1, and P1 stops meaning anything. A useful forcing function is a budget — no more than, say, twelve P1 rules for the whole service — which turns "should this page?" into a comparison against the existing twelve rather than a solo judgement.

Error budgets and burn-rate alerting

The most robust alerting scheme available ties the threshold to a business commitment rather than to a distribution. Start from an SLO: 99.9% of requests succeed over 30 days. That grants an error budget:

Text
30 days = 43,200 minutesbudget  = 43,200 x 0.001 = 43.2 minutes of failure per month

Define the burn rate as the observed error ratio divided by the budgeted one. Burning at 1x consumes the budget in exactly 30 days. Burning at 14.4x consumes it in 30/14.4 = 2.08 days. The alert then encodes both severity and urgency in one number, and the standard multi-window scheme is:

Burn rateLong windowShort windowBudget consumedAction
14.4x1 hour5 min2% in one hourPage
6x6 hours30 min5% in six hoursPage
3x1 day2 hours10% in a dayTicket
1x3 days6 hours10% in three daysTicket

Check the top row: one hour is 1/720 of a 30-day month, so at 14.4x you consume 14.4/720 = 2.0% of the budget in that hour. The short window is the trick that makes this responsive — the alert requires both windows to be burning, so it resolves quickly once the incident ends instead of staying lit for an hour after recovery.

Text
groups:  - name: slo    rules:      - alert: ErrorBudgetBurnFast        expr: |          (llm:error_ratio_1h  > 14.4 * 0.001)            and          (llm:error_ratio_5m  > 14.4 * 0.001)        labels:    {severity: page}        annotations:          summary: "Burning error budget 14.4x - 2% of the month in one hour"          runbook: "https://wiki.internal/runbooks/llm-error-budget"      - alert: LatencySLOBreach        expr: llm:latency_p95_5m > 2.0        for: 5m                          # five consecutive evaluations        labels:    {severity: page}        annotations:          summary: "p95 latency {{ $value | humanizeDuration }} exceeds 2s SLO"      - alert: InputDriftHigh        expr: model_feature_psi > 0.25        for: 2h                          # drift is slow; do not page on a blip        labels:    {severity: ticket}        annotations:          summary: "PSI {{ $value }} on {{ $labels.feature }} vs training baseline"

Routing: turning 40 alerts into one page

A single bad deploy will trip a dozen rules at once. Alertmanager's job is to make that one interruption rather than twelve, and three features do the work.

  • Grouping. Collapse alerts sharing labels — service, cluster, alertname — into one notification, held briefly so that related alerts arrive together rather than as a burst.
  • Inhibition. Suppress downstream alerts when an upstream one is firing. If ServiceDown is active, do not also page about HighLatency and ErrorRateHigh for the same service — they are consequences, and the responder already knows.
  • Silences. Time-bounded mutes with an author and a reason, applied during known maintenance. Bounded is the key word; a silence with no expiry is how an alert disappears permanently and nobody notices for a year.
Text
route:  group_by: [alertname, service]  group_wait: 30s              # let related alerts arrive before notifying  group_interval: 5m  repeat_interval: 4h  receiver: slack-default  routes:    - matchers: [severity="page"]      receiver: pagerduty      continue: true           # also post to Slack for visibility    - matchers: [severity="ticket"]      receiver: jirainhibit_rules:  - source_matchers: [alertname="ServiceDown"]    target_matchers: [severity=~"page|ticket"]    equal: [service]

Automated retraining, and the guardrails that make it safe

The tempting wiring is "drift alert fires → retraining pipeline starts → new model deploys". Every part of that chain has a specific way of going wrong.

TriggerSensibleDanger
Scheduled (weekly, monthly)Predictable; easy to reason about; catches slow driftRetrains when nothing changed, burning compute and review time
Performance-based (accuracy below threshold on labels)The most defensible trigger — it responds to the thing you care aboutNeeds labels, which arrive late; the damage is done by the time it fires
Drift-based (PSI above 0.25)Fast; fires before labels existThe most common cause of large PSI is an upstream data bug
Volume-based (50,000 new labelled examples)Ties retraining to genuinely new informationNew data may be biased towards whatever the current model surfaced

That third row is the one that turns an outage into a permanent regression. Recall the pattern: a partner starts sending distances in miles instead of kilometres, PSI spikes, the automated pipeline retrains on the corrupted feed, and the corruption is now in the weights. Fixing the pipeline no longer fixes the model. Drift opens a ticket. A human decides whether the answer is a retrain, a pipeline fix, or nothing.

Do the cost arithmetic too. Suppose a retrain costs 400 dollars of compute plus four engineer-hours of validation. A drift rule that produces one false trigger a week costs:

Text
52 x 400 dollars       = 20,800 per year in compute52 x 4 engineer-hours  = 208 hours per year of review

Five weeks of an engineer's year spent validating retrains that were never needed. Alert thresholds have a direct line to a budget.

The gate a challenger must pass

When a retrain does happen, "the new model is better" needs a number, because accuracy measured on a finite holdout is itself noisy. With a holdout of 2,000 examples and a champion accuracy of 0.91, the standard error of a single accuracy estimate is:

Text
SE = sqrt(0.91 x 0.09 / 2000) = sqrt(4.095e-5) = 0.0064 = 0.64 percentage points

So a challenger scoring 91.5% is +0.5 pp — comfortably inside the noise of a single measurement, and completely meaningless as evidence. A defensible gate requires the improvement to clear roughly two standard errors, so about +1.5 pp on this holdout, or better still uses a paired comparison (McNemar's test on the examples where the two models disagree), which is far more sensitive because it removes the variance shared by both models.

Python
def promotion_gate(champion, challenger, holdout, slices):    checks = {        # headline metric must beat the incumbent by more than measurement noise        "beats_champion": challenger.accuracy - champion.accuracy > 0.015,        # no slice may regress, however good the average looks        "no_slice_regression": all(            challenger.accuracy_on(s) >= champion.accuracy_on(s) - 0.01            for s in slices),        # the known-failure regression suite must not break        "regression_suite": challenger.score_on(REGRESSION_SET) >= 0.95,        # trained on data that passed validation        "training_data_clean": challenger.training_data.validation_passed,        # cost and latency did not silently blow out        "latency_ok": challenger.p95_latency_ms <= champion.p95_latency_ms * 1.10,    }    return all(checks.values()), checks

The no_slice_regression check earns its place repeatedly. A model can gain 2 points overall while losing 9 points on a minority segment, because the aggregate is dominated by the majority. Aggregate metrics hide exactly the failures that generate complaints.

Beyond the gate, the deployment itself should be staged: run the challenger in shadow first (it sees live traffic, its outputs are logged and never served), then canary to 5% of traffic with automatic rollback on any guardrail breach, then ramp. And keep the previous model artefact loadable, with a rollback that is one command and takes under five minutes. Automated promotion without automated rollback is not automation; it is a faster way to break production.

What this means when you build one

Three habits separate an alerting system that survives a year from one that becomes a muted channel.

Compute the false-alarm rate before you merge the rule. Number of checks per week times the tail probability, times the number of rules. If the total across your whole alert set is more than about five per week, you are building the 47-a-day channel and you will get there in about a quarter. This calculation takes two minutes and almost nobody does it.

Reach for persistence before you reach for a tighter threshold. At a matched false-alarm budget, a 2.33σ threshold requiring two consecutive breaches detects a 2σ regression in 10 checks where a 3.72σ single-breach rule takes 23. Then backtest both against 30 days of real history, because your metric is autocorrelated and the theoretical numbers are optimistic.

Review the alerts, not just the incidents. Once a month, list every alert that fired and mark each one: acted on, ignored, or auto-resolved. Anything in the "ignored" column for two consecutive months gets deleted or downgraded — no exceptions, no "but it might be useful someday". An alert nobody acts on is not free; it is spending the attention that the next real page will need.

The team from the opening did exactly this. They deleted 31 of their 40 rules outright, converted 6 into dashboard panels, rewrote the remaining 3 as burn-rate alerts with two-window conditions, and added a drift rule routed to a ticket queue rather than a pager. The following quarter their channel received 38 messages. Nine of them were real. The 02:14 alert would now be the only thing in the channel.