Mobile System Design Interview

Course Content

Mobile System Design Interview

11 sections · 23 lessons

Chat App: the outbox, ordering and follow-ups


A message the user pressed send on must never be lost and must never arrive twice. Those two requirements, together, are what the outbox exists for. The send path from the previous lesson wrote a message row and an outbox entry in one transaction; this lesson designs what happens to that entry afterwards.

It then takes the second hard problem in chat — getting one stable order out of two devices with two clocks — and finishes with the five follow-ups this problem reliably attracts.

Why an in-memory retry queue fails

The naive design keeps unsent messages in a list and retries them. Two failures, both common:

The process dies. The user sends a message on a train, the screen locks, the operating system reclaims the app. The queue was in memory. The message is gone, and the sender saw it appear on screen, so they believe it was sent.

The response is lost. The request reached the server and the reply did not — the tunnel, the tower handoff. The client retries. Without a way to recognise the repeat, the recipient sees the message twice.

The outbox design

Persist the queue. The outbox is a table, written in the same transaction as the message row. It survives process death by construction.

Generate the identifier on the client. The message carries a client-generated unique identifier from the moment it is created. The server treats it as an idempotency key: a second send with an identifier it has already stored returns the original result rather than creating a new message. That single decision makes retrying safe, and it must be requested from the server team explicitly.

Drain with backoff. A worker takes outbox entries in order, sends them, and on success updates the message row and deletes the entry. On a retryable failure it backs off. On a permanent failure — the conversation was deleted, the user was blocked — it marks the message failed and stops.

Preserve order within a conversation. Drain one conversation's outbox serially. Sending three queued messages in parallel means they can arrive in any order, and the recipient sees a jumbled exchange.

The state machine

Each message the user sends moves through a small set of delivery states, with an explicit failed state the user can retry.

The outboxmsg-1 · sent, ackedmsg-2 · sendingmsg-3 · queuedmsg-4 · queuedwritten to disk before the UI shows the message — so a crash mid-send losesnothingCOMPOSINGQUEUEDSENTDELIVEREDREADFAILEDsend tappedserver ackedrecipient device ackedrecipient opened itretries exhausteduser taps retryThe message id is generated on the client, so a retry after a lost response is idempotent and the recipient never sees the message twice.
QUEUED is the state that makes the app feel instant: the message is durable and visible before the network is involved at all.

Reading the state machine honestly

Two of these transitions depend on the recipient's device, not the server: delivered when their device receives it, read when they open the conversation. Both require the recipient to be online, so a message can sit at sent for days. The user interface must not imply failure — a single tick means "the server has it", which is the promise you can actually keep.

Offline: the outbox fills, every message shows a clock, and the count of pending messages is visible. On reconnect the outbox drains before the resync, so the user's own messages appear in the right place.

Ordering and sync

The outbox guarantees each message arrives exactly once. It says nothing about where it appears in the list. Two devices, two clocks, one conversation: getting a stable order out of that is harder than it looks, and the wrong answer is the intuitive one.

Server sequence numbers, and the gap they exposeseq 101seq 102seq 105seq 106nulllastcontiguousgapdetected103 and 104 never arrived, so the client backfills that range before rendering past the gap.
Device clocks disagree by minutes, so only a server-assigned sequence can order a conversation.

Why device clocks cannot order messages

The intuitive design stamps each message with the sender's local time and sorts by it. It fails in ways users notice:

  • Phone clocks drift, and users set them manually. A device five minutes fast puts every message it sends above messages sent after it.
  • Time zone and daylight-saving changes can move a clock backwards.
  • Two messages sent in the same millisecond have no defined order, so different devices sort them differently and two people in the same conversation see different transcripts.

That last one is the killer: order must be identical on every device, and a value each device computes for itself cannot guarantee that.

Server sequence numbers

The server assigns a monotonically increasing sequence number per conversation as it accepts each message. That number is the sort key. It is assigned in one place, so every device sorts identically, and it is dense enough to detect gaps.

The client keeps last_seq per conversation — the highest contiguous sequence it holds.

The provisional-order problem

A message the user sends offline has no sequence number yet, because no server has seen it. It still has to appear somewhere in the list.

The usual solution is a composite sort key: order by (seq, local_counter) where unsent messages carry a sequence higher than anything received — so they sit at the bottom, where the sender expects them. When the server acknowledges and assigns a real sequence number, the row is updated and may move. Moving it a short distance is acceptable; the alternative, holding the message invisible until acknowledged, is not.

Gap detection and backfill

Messages arrive out of order or go missing. The socket delivers 101, 102, then 105. The client holds 103 and 104 nowhere.

Text
onMessage(m):    if m.seq == last_seq + 1:        // in order        insert; last_seq = m.seq        drain any buffered messages that now connect    else if m.seq > last_seq + 1:    // gap        insert into a holding area        fetch(conversation, from: last_seq + 1, to: m.seq - 1)    else:                            // already have it        ignore  (duplicate — normal after a reconnect)

Two properties worth stating: duplicates are expected and discarded silently by primary key, and a gap triggers a bounded fetch rather than a full resync.

Coming back after a long time offline

A device offline for two weeks may be tens of thousands of messages behind. Requesting all of them at once is a large download on a metered connection and a memory risk.

The right shape: fetch the most recent page per conversation so the user can read immediately, then backfill older ranges lazily as they scroll, marking the unfilled region so the conversation shows "loading earlier messages" rather than pretending it is complete. If the gap is beyond the server's retention for delta queries, fall back to a bounded history fetch and accept that very old messages come from history pagination instead.

Follow-ups

Five questions this problem reliably attracts.

Letting history grow against pruning itKeep everything on device• Instant scroll into old threads• Database grows without bound• Cold start and queries slow downPrune to a window• Keep recent messages per thread• Older pages refetched on demand• Bounded storage, predictable start-up
Pruning per conversation rather than globally keeps active threads fast and dormant ones cheap.

Media attachments with progress and resumption

An attachment is two operations: upload the bytes, then send a message referencing them. Doing it in that order means the message row exists locally with a local file path and a pending state, the upload runs as a resumable chunked transfer with progress written to the database (so the progress bar survives the screen being closed), and the send is enqueued only once the upload completes.

Hand the transfer to the platform's background transfer mechanism, so it continues after the app is suspended. Show the local file immediately — the sender sees their own photo without waiting for a round trip.

Pagination over a long history

The conversation screen queries the most recent 50 messages and pages backwards as the user scrolls up. Reverse-cursor pagination against the local database, and when the local history runs out, a network fetch for older ranges that writes into the same table. The user cannot tell where the boundary is, which is the point.

Database growth and pruning

Text is cheap; media is not. A defensible policy: keep all message text indefinitely, prune downloaded media not opened in 30 days when total media exceeds a ceiling, and always keep a thumbnail so the conversation still shows something with a tap-to-redownload affordance. Give the user a per-conversation storage view and a clear button — being able to see it is part of the design.

Typing indicators and their cost

The naive implementation sends an event per keystroke. At 4 characters per second that is 4 messages per second, each a radio transaction, for information worth almost nothing.

The design: send "typing" at most once every 3–5 seconds while typing continues, expire it client-side after ~6 seconds without a refresh (so a crashed sender does not leave a stuck indicator), and send nothing at all when the connection is metered and the setting allows it. Never persist typing state to the database — it is ephemeral UI state, tier one from Step 4: data flow and state.

What end-to-end encryption changes on the client

More than candidates expect, and all of it client-side:

  • Key management becomes an app concern: generating, storing in the secure enclave, publishing public keys, and re-establishing sessions per conversation.
  • Multi-device stops being free. Each device has its own keys, so a message must be encrypted once per recipient device.
  • Server-side search and history restore disappear. The server cannot read messages, so search must run over the local database, and a new device cannot be handed history by the server.
  • Backup becomes a designed feature — an encrypted export with a user-held key — rather than something the server does invisibly.
  • Push payloads carry no content, so the notification must say "New message" until the app decrypts locally, or the notification handler must be able to decrypt.