Course Content
System Design Interview
31 sections · 71 lessons
Notification System: scope, scale and high-level design
The prompt: "Design a notification system."
Before you read another line, put a timer on for 45 minutes and attempt it on paper. Draw the boxes, write the API, do the arithmetic. Reading this lesson without having struggled with the problem first produces recognition, not skill.
This lesson takes the problem from a vague prompt to a first architecture that survives. It narrows the prompt with five questions, puts numbers on the load — including the one number that actually sizes the system — and then builds the smallest design that could work, breaks it, and replaces it with the shape that production systems use.
Why this prompt is vague on purpose
"Notification" means at least four different products. A push message on a phone, a text message, an email, and a red badge inside the app share a name and share almost nothing else: different providers, different latency expectations, different failure modes, and different legal obligations.
The interviewer knows this. The first ten minutes exist so you can narrow it.
The five questions that shape everything after
1. Which channels are in scope? Mobile push, SMS (short message service, a text message sent over the mobile network), email, and in-app inbox are the usual four. Each has a different third-party provider and a different delivery guarantee. Ask, then scope: "I'll design for push, SMS, and email, with in-app as a fourth channel that we own end to end."
2. Real-time or best-effort? A two-factor authentication code that arrives in 90 seconds is a failed product. A weekly digest email that arrives 20 minutes late is fine. If the answer is "both", you have a priority requirement, and priority requirements always mean separate queues later.
3. Can users opt out, and at what granularity? Per channel, per category, or a single global switch? Opt-out is not a feature bolted on at the end — it decides where the preference check sits in the pipeline, and getting that wrong is a compliance problem, not a bug.
4. Are notifications event-triggered, scheduled, or broadcast? Event-triggered ("your order shipped") is one recipient. Broadcast ("the sale starts now") is every recipient at once, and the burst it creates is usually the hardest part of the design.
5. Who calls this system? If the answer is "forty internal services", then the system is a platform with an API, templates, and a contract — not a function inside one service.
What you are not designing
Say what you are excluding, out loud. Notification content generation, the machine-learning model that decides whether a notification is worth sending, and the analytics warehouse are all separate systems. Scoping them out buys you the time to go deep on delivery, which is where the interesting engineering is.
Requirements and scale
Numbers first, architecture second. The arithmetic here is what makes the queue in the high-level design below an obligation rather than a preference.
Functional requirements
- Any internal service can request a notification for a user, in a named category.
- The system resolves the user to their devices, addresses, and phone numbers.
- It honours the user's preferences and quiet hours.
- It renders a template into the right language and channel format.
- It delivers through the right provider, retrying on failure.
- It records what was sent, what was delivered, and what failed.
Non-functional requirements
- At-least-once delivery, with no duplicate visible to the user. Reliability: exactly-once as a goal is entirely about how those two survive together.
- Transactional notifications in under 5 seconds end to end, at the 95th percentile. Marketing may take minutes.
- The system never loses a request it has acknowledged. If we return 202 Accepted, it gets sent or it lands in a dead-letter queue for a human.
- Provider failure is isolated. An email provider outage must not delay push.
The estimate
All figures below are invented for a plausible mid-size consumer product, and every one is computed rather than quoted. Following Section 3, a day is rounded to 100,000 seconds.
| Input | Assumption |
|---|---|
| Registered users | 100 million |
| Daily active users | 20 million |
| Push per active user per day | 5 |
| Emails per active user per day | 1 |
| SMS per active user per day | 0.1 |
- Push: 20M × 5 = 100M/day → 100M ÷ 100,000 = 1,000 push/s average. Evening peak at 5× average = 5,000/s.
- Email: 20M × 1 = 20M/day → 200/s average, 1,000/s peak.
- SMS: 2M/day → 20/s average. Low volume, highest cost per message.
The number that actually sizes the system
A marketing broadcast to all 20 million active users, sent over a 10-minute window:
20,000,000 ÷ 600 seconds = 33,000 messages per second
That is 33× the steady-state push rate, from a single API call. And a typical push provider contract might allow, say, 10,000 messages per second per account. The burst exceeds what the provider will accept by more than 3×.
Storage
Device tokens: 100M users × 1.5 devices × ~200 bytes = 30 GB. Small enough to sit in a single replicated database, and a strong hint that the token store is not where the difficulty lives.
A delivery log of every attempt, at 120M notifications/day × ~300 bytes, is 36 GB/day, or about 13 TB/year. That belongs in a time-series or columnar store with a retention window, not in the operational database.
The high-level design
With the numbers in hand, start with the smallest thing that could work, then break it.
Version one, and why it fails
The order service wants to tell a user their parcel shipped. It calls Apple Push Notification service directly.
Five things go wrong, and each names a component:
- Every service reimplements it. Forty services, forty copies of retry logic, forty places the preference check can be forgotten.
- The provider call is in the request path. A 2-second provider timeout becomes 2 seconds added to placing an order.
- A provider outage loses notifications. There is nowhere to park a message that cannot be sent right now.
- Bursts are passed straight through. The broadcast estimated above hits the provider at 33,000/s and gets rejected.
- No preference enforcement. A user who opted out gets messaged anyway.
Version two: the shape that survives
Three layers, with a buffer between them.
Notice where the rate changes: producers write at burst speed, workers read at whatever the provider accepts, and the log holds the difference.
Why the queue is not optional
Three separate jobs, and no simpler component does all three:
- Rate shaping. Producers write at 33,000/s; workers read at whatever rate the provider accepts. The log holds the difference. A 10-minute backlog at 23,000/s of overflow is 13.8 million messages — at 500 bytes each, about 7 GB of buffered data. That is comfortable for a distributed log and impossible for an in-memory buffer.
- Retry without blocking. A failed send is a message that stays available, not a caller who waits.
- Isolation. One topic per channel means an email backlog does not slow push.
Say "a distributed log" for the concept; Apache Kafka is the common implementation, and a managed cloud equivalent is a fine substitute. Message queues and event streaming covers the queue-versus-log distinction, and Section 21 builds one.
The API
POST /v1/notifications{ "recipient": {"user_id": "u_4412"}, "category": "order_shipped", "template_id": "order_shipped_v3", "params": {"order_id": "8891", "eta": "2026-09-02"}, "channels": ["push", "email"], "priority": "transactional", "dedupe_key": "order-8891-shipped"}→ 202 Accepted {"request_id": "nr_01H..."}Two decisions worth defending. It returns 202 Accepted, not 200 OK, because acceptance is not delivery and the API should not lie about that. And dedupe_key is supplied by the caller, which Reliability: exactly-once as a goal explains.
The data model
usersanddevices— device token, platform, last seen. Relational; small.preferences— user, category, channel, allowed. Read on every send, so cached.templates— template ID, locale, channel, body.delivery_log— append-only, one row per attempt, in a time-series store.