AI in the SOC: why precision beats recall in a flood of alerts

JR

Jai Rao

August 22, 202617 min read

A defender-side look at machine learning in security operations: the base-rate arithmetic that buries analysts, and where ML actually reduces the load.


A detector advertised at 99.9% accuracy sounds finished. Point it at a real enterprise event stream and it will produce something like ten thousand alerts a day, of which roughly ninety will matter. Nobody triages ten thousand alerts a day. The team stops reading the queue, the tuning backlog grows, and six weeks later the detector is switched off in all but name — not because it was inaccurate, but because accuracy was never the quantity that decided whether it was usable.

That gap is the first thing to understand before pointing machine learning at security operations, and it has almost nothing to do with models. It is arithmetic about rarity. Attacks are vanishingly uncommon relative to ordinary computing activity, and that rarity dominates every performance figure a demo will show you. Everything else here — anomaly detection, alert clustering, model-assisted triage, generated rules — either respects that arithmetic or loses to it.

The arithmetic that buries a SOC

Take a mid-sized enterprise: a few thousand endpoints, a handful of cloud accounts, ordinary business software. Assume the telemetry reaching the SOC — process executions, authentication events, DNS queries, flow records, cloud API calls — comes to ten million records a day; adjust for your own environment and the conclusion holds. Now assume an intrusion is genuinely under way, producing one hundred individually detectable malicious events that day. Prevalence of "malicious" in the stream is 100 in 10,000,000: one in a hundred thousand.

Give the detector a recall of 0.90 — it catches nine of every ten malicious events, which is better than most real detections manage — and a false-positive rate of 0.1%, meaning one benign event in a thousand is misclassified. Both numbers sound respectable. This is what they produce, at that prevalence and at three progressively stricter false-positive rates.

Text
def precision(prevalence, recall, fpr):    tp = recall * prevalence    fp = fpr * (1 - prevalence)    return tp / (tp + fp) if tp + fp else 0.0EVENTS = 10_000_000MALICIOUS = 100prev = MALICIOUS / EVENTS            # 1 in 100,000for fpr in (1e-3, 1e-4, 1e-5, 1e-6):    p = precision(prev, 0.90, fpr)    alerts = round(0.90 * MALICIOUS + fpr * (EVENTS - MALICIOUS))    print(f"FPR {fpr:9.6f}  alerts/day {alerts:6,}  precision {p:6.2%}")# FPR  0.001000  alerts/day 10,090  precision  0.89%# FPR  0.000100  alerts/day  1,090  precision  8.26%# FPR  0.000010  alerts/day    190  precision 47.37%# FPR  0.000001  alerts/day    100  precision 90.00%

The first row is the "99.9% accurate" detector: fewer than one alert in a hundred is real. Improve it tenfold and one in twelve is real — still a queue that trains people to close things unread. Precision only becomes tolerable in the third row, where the false-positive rate has fallen to one in a hundred thousand and finally matches the prevalence of what is being detected. That is the rule worth memorising: a detector's false-positive rate must fall to roughly the base rate of the attack before half its alerts are true.

Convert those rows into staffing and the point lands harder. At ten minutes of triage per alert, the first row is about 1,680 analyst-hours a day — two hundred eight-hour shifts, for a team that probably has five people. The second row is 180 hours, or twenty-odd full-time analysts doing nothing else. The third row is roughly 32 hours: four analysts, with about half of what they open turning out to be genuine. Only the third row describes a detection programme that exists. The others describe a dashboard.

In plain text, precision equals (recall x prevalence) divided by (recall x prevalence + FPR x (1 - prevalence)). Prevalence is tiny, so (1 - prevalence) is effectively 1 and the denominator is dominated by FPR alone. Alert volume is set almost entirely by false-positive rate multiplied by event volume. Recall barely enters it.

Precision is the number that pays for itself

Most ML practice optimises for recall, and in security that instinct is usually wrong at the event level. Compare two improvements against the third row above. Push recall from 0.90 to 0.99 and you gain nine detected events; precision moves from 47% to 50%. Cut the false-positive rate tenfold instead and precision goes to 90% while the queue shrinks by half. The second change gives your analysts their afternoons back. The first is a rounding error on a graph.

The honest objection is that recall is the whole reason a SOC exists; a missed intrusion is the failure you actually fear. The resolution is that recall matters at the level of the campaign, not the individual event. An intrusion is not one event — it is initial access, persistence, credential use, lateral movement and collection, spread over hours or weeks and leaving traces in several telemetry sources. If a campaign produces twenty observable events and per-event recall is only 0.30, the chance of catching it at least once is 1 - 0.7 raised to the 20th power, about 99.9%. Real events are not independent, so the true figure is lower — but the direction holds, and it is why you can trade event-level recall for precision without losing campaign coverage.

Precision also compounds. A queue where half the alerts are real gets read carefully, so dispositions are recorded accurately and you accumulate labels worth training on. A queue where one in a hundred is real gets closed in bulk with the default disposition. Low precision does not merely waste time; it destroys the data you would need to fix it.

Labelled attacks are scarce, and the labels you have are skewed

The supervised-versus-unsupervised question in security is usually presented as a modelling choice. It is really a question about where labels come from, and the answers are uncomfortable. Analyst dispositions only cover events that some existing rule already fired on. Red-team exercises reflect one team's playbook and schedule. Public malware corpora contain samples that were caught and published — by construction, the ones that evaded detection are absent. Threat-intel indicators are stale the moment they are shared.

Put plainly: you have labels for the attacks you detected. The attacks you missed sit in your data unlabelled, indistinguishable from benign traffic, and they are precisely the failure modes you most want to learn. That survivorship bias is not fixed by more data or a bigger model.

Class imbalance compounds it. At one malicious event in a hundred thousand, a model that always answers "benign" scores 99.999% accuracy. Class weights, focal loss and resampling change what the optimiser cares about without adding information the training set never had: they make a model willing to guess "malicious", not right about it.

ApproachWhat it needsWhere it holds upHow it fails
Supervised classificationMany labelled examples of both classesPhishing email and URL scoring, malware file triage, algorithmically generated domain names — stable definitions, cheap labels, both classes plentifulCovers only attacks you already catch; a novel technique is a silent miss with a confident score
Unsupervised / anomaly scoringA baseline of normal, no labelsBeaconing regularity, rare parent-child process pairs, first-seen destinations, unusual authentication geometryBenign novelty is abundant, so raw anomaly volume is high and severity is unknowable
Rules and invariantsAn analyst who can describe the behaviourPrecise, explainable, auditable, cheap to run and cheap to testBrittle, and blind to anything nobody thought to write down

So supervised learning belongs on narrow sub-problems with cheap labels and a stable class definition — phishing scoring behaves far more like spam filtering than like intrusion detection. Unsupervised methods belong upstream of the queue, producing rankings rather than verdicts. Neither replaces rules, which are the only artefact that encodes "this must never happen here" in an auditable form.

What unsupervised models are genuinely good at

Two patterns show why label-free methods earn their place. The first is beaconing. Implants that need instructions tend to check in on a schedule; people do not. Human-driven traffic is bursty and irregular, machine-driven traffic is metronomic, and the regularity of the gaps between connections is measurable without a single label. A coefficient of variation over inter-arrival times captures most of it.

Text
import statisticsdef beacon_score(timestamps):    """Regular gaps look scheduled; human traffic is bursty and irregular."""    if len(timestamps) < 12:        return 0.0    gaps = [b - a for a, b in zip(timestamps, timestamps[1:])]    mean = statistics.fmean(gaps)    if mean <= 0:        return 0.0    cv = statistics.pstdev(gaps) / mean      # coefficient of variation    return max(0.0, 1.0 - cv)human = [0, 4, 61, 63, 190, 900, 905, 1400, 1402, 1500, 3000, 3100]scheduled = [i * 300 + (i % 3) for i in range(12)]print(round(beacon_score(human), 2), round(beacon_score(scheduled), 2))

That prints 0.0 1.0: the browsing pattern scores nothing, the five-minute schedule scores the maximum. It also scores every telemetry agent, NTP client and software updater in your fleet just as highly — the base-rate problem returning through the side door. The fix is not a better score but a conjunction. Join it with destination rarity: a perfect score to an endpoint four thousand hosts also contact is your patch server; a perfect score to one that exactly one host has ever contacted deserves attention.

The second pattern is impossible travel — one identity authenticating from two places further apart than the elapsed time allows. No labels required, just clocks and geometry. Alone it is low-precision: VPN egress, corporate proxies, wrong mobile-carrier IP geolocation, and a phone syncing mail on cellular beside a laptop on office wifi all generate it constantly. It works as a feature rather than an alert — impossible travel plus a first-seen device plus a changed MFA method has a far lower base rate than any of its parts.

Conjunctions are the most reliable precision lever available and they need no model at all. Two signals at one-in-a-thousand each would, if independent, combine to one in a million — the gap between the first and last rows above. Real signals correlate, so measure the conjunction on historical data rather than multiplying and hoping.

Collapsing the flood before a human sees it

An alert flood is rarely ten thousand distinct problems. It is usually three problems reported ten thousand times: a scanner nobody allowlisted, a rule interacting badly with one build server, a backup job touching files in a pattern the model finds strange. Before reaching for a better detector, build a case key — the tuple capturing what an analyst would call "the same thing happening again".

Text
from collections import defaultdictdef case_key(a):    """What an analyst would call 'the same thing happening again'."""    return (a["rule"], a["host"], a["dest_asn"])alerts = (    [{"rule": "R-104", "host": "web-07", "dest_asn": "AS15169"}] * 812    + [{"rule": "R-104", "host": "web-08", "dest_asn": "AS15169"}] * 640    + [{"rule": "R-221", "host": "hr-lap-3", "dest_asn": "AS9009"}] * 3)cases = defaultdict(int)for a in alerts:    cases[case_key(a)] += 1print(len(alerts), "alerts ->", len(cases), "cases")for key, n in sorted(cases.items(), key=lambda kv: kv[1]):    print(f"  {n:5} x {key}")

1,455 alerts become 3 cases, and the sort is deliberately ascending so the three-event case surfaces above the two noisy ones. That ordering matters: the loud groups are almost always misconfiguration, and the rare singleton is where a real intrusion tends to appear. Sorting a case list by volume, descending, is the most common way teams hide the thing they were looking for.

Beyond exact keys, embedding alert text and clustering the vectors catches near-duplicates that different tools word differently — worth doing once several products write into one queue. Two constraints hold either way. Suppressed alerts must be counted, never silently dropped: "812 occurrences since Tuesday" is a different investigation from "1 occurrence". And grouping improves the presentation of a low-precision detector without making it precise. If your key buckets a real intrusion in with a benign flood, you have made things worse.

Where language models actually buy time

The expensive part of triage is not the decision. It is assembling the context that makes a decision possible: who owns this host, whether the binary is signed and how common it is across the fleet, whether this identity has used this device before, what the same rule did last month, whether an open change ticket explains it. That is ten to twenty minutes of tab-switching per alert, nearly all of it deterministic lookup.

This is where a language model genuinely helps. It can drive those lookups as tool calls into asset inventory, identity, endpoint telemetry and ticketing, then write a short paragraph of what it found with the evidence attached. It is not judging maliciousness; it is making evidence legible so a human judges faster. Build it so every claim carries an identifier linking back to the record behind it. A fluent summary an analyst cannot verify is worse than none, because confident phrasing borrows credibility the model has not earned.

One design property deserves explicit attention: in a SOC the input is attacker-controlled by definition. Filenames, user agents, email subject lines, HTTP paths, certificate common names and process command lines are all fields an adversary writes into, and all of them land inside your summarisation prompt. Treat every retrieved field as untrusted data, never as instruction. The mitigation is architectural rather than textual — keep the model that reads attacker-supplied content separate from the tools that can act. If the summariser's only output is text in a case note, injected instructions are an annoyance. If it also holds a token that can isolate hosts, they are a control channel.

Drafting detection rules a human then owns

A model that has your log schema in context can turn "alert when a service account is used for interactive logon" into syntactically valid detection content in seconds. Syntax is the tedious part, and this is a good use of generation.

Text
title: Service account used for interactive logonstatus: experimentaldescription: Service accounts should authenticate non-interactively.  An interactive session for one deserves an analyst's attention.logsource:  product: windows  service: securitydetection:  logon:    EventID: 4624    LogonType: [2, 10]        # console, remote interactive  svc_naming:    TargetUserName|startswith: 'svc-'  filter_break_glass:    TargetUserName: 'svc-emergency-admin'  condition: logon and svc_naming and not filter_break_glassfalsepositives:  - Engineer debugging a service by logging in as its account  - Vendor installer running under a service identitylevel: medium

What the model cannot supply is the part that determines whether this rule is usable: the filter_break_glass exclusion and the false-positive list. Those come from knowing that one account is legitimately used by hand during incidents, and that your patching vendor does something odd on Tuesdays. That knowledge lives in your environment, not in the model's training data.

So the pipeline is generation, then measurement, then review. Backtest every candidate against thirty to ninety days of historical telemetry and count what it would have produced. A rule that would have fired four thousand times last month is not a detection, it is a self-inflicted incident. That backtest is the false-positive rate from the first section, measured in your environment instead of assumed. Keep rules in version control, review them as code, and put the historical hit count in the pull request.

Detectors rot, because the adversary reads your output too

Two kinds of drift act on security models, and they need different responses. Benign drift is the fleet changing underneath the baseline: a new endpoint agent rolls out, a team adopts a SaaS tool, a data centre migration reroutes traffic, an office reopens. This is most of your drift, it is constant, and it shows up as an anomaly detector suddenly firing on the new normal.

Adversarial drift is different in kind. The distribution moves because someone is deliberately moving away from your decision boundary. A recommender's drift is incidental; this one is directed, and it responds to what you leak — including alert text visible in a console an intruder has reached. So a detector's measured precision has a shelf life. Re-measure on recent data rather than the holdout set from the week you shipped, and watch per-rule volume in both directions: a rule that goes unexpectedly quiet is as interesting as one that goes loud.

This is also the argument against putting everything into learned models. Where a behaviour can be stated as an invariant — no production database reachable from a guest network, no service account holding interactive session rights — write it as policy and enforce it. Policies do not drift, because they are not statistical estimates. Learned models cover the space you cannot articulate; invariants cover the space you can. Wire the two differently: when a ranking model drifts your priority order degrades, but when a blocking control drifts you have caused an outage.

Which actions still need a human hand

Isolating a host, disabling an account, blocking an egress range, killing a process tree — each is reversible in principle and each can break something that matters. Auto-isolating a domain controller because a model scored 0.94 is an outage you inflicted on yourself, and it will end the automation programme faster than any missed detection.

Gate automation on reversibility and blast radius, not on confidence score. Cheap, local, reversible collection can run unattended — pull a file hash, snapshot a VM, capture a process tree, gather recent authentication history for one identity — because the worst case is wasted compute. Anything reaching past a single user or host, such as network-level blocks or credential revocation for shared identities, needs a person to approve it.

Confidence thresholds make a poor gate on their own, and the reason is worth stating precisely. A score is calibrated against the distribution the model was trained on, and the incidents where you would most value autonomous action are the novel ones furthest from that distribution. The model's confidence is least trustworthy exactly where the stakes are highest.

Make approval fast rather than optional. Pre-stage the action, state the predicted blast radius concretely — "disconnects 1 host, terminates 3 active sessions, stops the named service" — then make it one click with a one-click undo. Approval latency is a real cost during an incident; the answer is better approval design, not removing the human. The approve-or-reject decision is also the highest-quality label your organisation produces, being the only one made with the full picture in view. Log it as training data.

Judging the programme, not the model

Model metrics on a holdout set tell you very little about whether a detection programme works. Measure the queue instead. Per rule and per model: alerts per day, cases per day after grouping, precision estimated by sampling and auditing dispositions rather than trusting free-text closures, median time to disposition, and the share of the queue that is never opened at all.

That last number is the one most teams do not track and the one that decides everything. If a third of your alerts age out unread, your effective recall is your model's recall multiplied by the probability a human ever looks — so a detector with 0.95 recall sitting in an unread queue performs worse than one with 0.60 recall in a queue that gets fully worked. Alert fatigue is not a morale problem to be managed with better dashboards. It is a term in your detection coverage, and precision is the lever that moves it.

The build order follows. Instrument the queue first; you cannot improve a volume nobody counts. Add case grouping second — cheap, and often more analyst time reclaimed than any model will manage. Add enrichment and summarisation third, since it attacks the largest fixed cost per alert. Then conjunction rules and invariants, which buy precision without training anything. Reach for learned detectors last, starting where labels are cheap and plentiful.

Four signs tell you it is working, none of them an F1 score. The queue is read to the bottom every shift. Precision is a measured number somebody owns, not a claim on a slide. Every rule change arrives as a reviewed diff with a backtest attached. And when the system is wrong, the analyst's correction lands somewhere that changes the next version rather than in a closed ticket nobody reads. Get those four in place and the models become straightforward to improve. Skip them and no model will save you, because the constraint was never the model.