System Design Interview

Course Content

System Design Interview

31 sections · 71 lessons

Email Service: delivery reliability, spam and follow-ups


The storage and search design from the previous lesson covers mail that has already arrived. This lesson covers the two directions where other servers decide the outcome: sending mail out, and keeping unwanted mail from coming in. It ends with the follow-up questions interviewers use to probe the rest of the design.

Outbound mail is the part of the system where you depend entirely on servers that do not answer to you.

An outbound message that is not acceptedSend attempt4xx: tryagain laterBackoff andretry queue5xx: bounceto senderReputationadjustsGreylisting rejects the first attempt on purpose; only real senders return.
A temporary rejection is the normal case, so the retry schedule is a feature of the system rather than an error path.

Queue, retry, and backoff

Every outbound message goes into a queue with per-destination-domain partitioning, so a slow or broken destination cannot block delivery to everyone else. Delivery workers pull, attempt SMTP, and interpret the response code.

Response codes are a three-way branch. A 2xx means delivered and the message is done. A 5xx is permanent — no such mailbox, message rejected — so stop and bounce. A 4xx is temporary — try again later.

Retry schedule for 4xx: exponential backoff with jitter, over a long horizon. A typical sequence is a few minutes, then tens of minutes, then hours, giving up after somewhere around 24 to 72 hours. The jitter matters because a receiving server that came back after an outage should not be hit by every queued message at the same instant — the same thundering-herd problem as a cache stampede.

Bounce handling. Giving up produces a bounce message back to the original sender describing the failure. Bounces must not themselves bounce, which is why they are sent with an empty envelope sender — a detail that exists specifically to break the loop that would otherwise form.

Greylisting, and why it exists

A receiving server temporarily rejects mail from a sender it has never seen, expecting a legitimate server to retry a few minutes later and a spam tool to move on. Retrying correctly is therefore a deliverability requirement, not only a reliability one: a sender that treats a 4xx as permanent will have its mail silently dropped by every greylisting recipient.

The visible cost is that first-contact mail can be delayed by minutes. Worth knowing, because "why did the first email to this domain take eight minutes" is a real support question with this answer.

Reputation: the part engineers underestimate

Whether your mail reaches an inbox, a spam folder, or nowhere is decided mostly by the receiving side's assessment of your sending reputation. The mechanisms:

  • SPF — Sender Policy Framework — a Domain Name System record listing which servers may send for a domain.
  • DKIM — DomainKeys Identified Mail — a cryptographic signature over the message, verified against a public key in the sender's Domain Name System records.
  • DMARC — Domain-based Message Authentication, Reporting and Conformance — a published policy saying what to do when SPF and DKIM fail, plus a reporting address.

Beyond authentication, receiving providers score sending internet-protocol addresses and domains on complaint rates, bounce rates, spam-trap hits, and volume patterns. The architectural consequences are concrete: separate sending pools so transactional mail is not poisoned by marketing mail, warm-up ramps for new addresses because sudden volume from an unknown sender looks like spam, and feedback loop processing so complaints suppress future sends to that recipient automatically.

Spam filtering, as a layered system

Now the inbound direction. Filtering is cheapest at the earliest layer, so order the layers by cost.

Filter as early as it is cheapConnection: IP reputationEnvelope: SPF and DKIMContent: a classifierUser: report and learn
Each layer costs an order of magnitude more than the one above it, so ordering them by cost is the whole design.
  1. Connection-level. Reject connections from addresses on reputation blocklists before any message data is transferred. Cheapest possible rejection.
  2. Envelope-level. Reject unknown recipients, enforce per-sender rate limits, apply greylisting to unknown senders.
  3. Content-level. Score the message body and headers with a classifier. Historically naive Bayes over token frequencies; in practice a blend of models and rules, retrained continuously because spam adapts adversarially.
  4. Post-delivery. Users mark messages as spam, which feeds training data and reputation scores. Retroactively moving already-delivered messages out of inboxes when a campaign is identified late is a real and useful capability.

The design consequence of the adversarial nature is that filtering must be retrainable on a short cycle, with a feedback loop from user reports. A static filter degrades within weeks.

Threading and conversation view

Threading is a graph problem over headers. Each message carries a Message-ID, and replies carry In-Reply-To and References naming their ancestors. Build the tree by linking those identifiers.

It breaks constantly, because some clients omit the headers and some strip them. The fallback is heuristic: group by normalised subject — stripping "Re:" and "Fwd:" prefixes — plus the participant set, plus a time proximity window. Mention both the correct mechanism and the heuristic fallback, because knowing the fallback exists is what shows familiarity with the real problem.

Store the thread identifier on the metadata row at delivery time, so listing a conversation is a lookup rather than a graph traversal per request.

Sync across devices

Three devices must agree on what is read, deleted, and labelled. The mechanism is a per-user monotonically increasing change identifier: every mutation to a mailbox increments it and is recorded in a change log. A client stores the highest identifier it has seen and asks "what changed since N?", receiving a compact delta rather than a full refresh.

That single mechanism gives you offline support, cheap reconnection, and multi-device consistency, and it is the same idea as the sync design in Google Drive's sync architecture.

Storage tiering

Mail older than a year is read very rarely. Moving bodies and attachments — never metadata — through hot, warm, and cold tiers is the largest cost lever in the system.

The arithmetic. Of 200 PB, suppose 10% is under 90 days old and actively read, 15% is between 90 days and a year, and 75% is older. Cold object storage commonly costs on the order of a quarter of standard storage per byte, with retrieval fees and higher latency. Moving 150 PB from standard to cold saves roughly 75% of the storage cost on 75% of the data — call it half the total storage bill. Exact prices vary by provider and change, so present it as a ratio rather than a currency figure.

The cost is retrieval latency on old mail: seconds instead of milliseconds. Acceptable, and the user interface should say "retrieving" rather than appear broken.