Course Content
Mobile System Design Interview
11 sections · 23 lessons
Chat App: requirements, transports and client architecture
Prompt: "Design a chat application."
This is the same product as Design a Chat System in System Design Interview, seen from the other side of the wire, and the contrast is worth holding in your head as you read. That section designs the message store, the fan-out, and the presence service. This one assumes all of that works and designs the app: the socket, the local database, the outbox, and what a user sees when their train enters a tunnel mid-sentence.
If you have read both, you have the whole system. If you are interviewing for a mobile role, this side is the one you will be scored on. This lesson takes the round up to the architecture; the next covers the outbox, ordering and the follow-ups.
The questions worth asking
"One-to-one, group, or both?" Take both, and note that groups make ordering harder because more writers contend for the same conversation. Cap group size in the requirements — "up to 256 members" — so you are not accidentally designing a broadcast system.
"How much history is available on the device?" This decides your storage budget and your sync design. The useful answer is "recent history locally, older history fetched on demand", which is what real clients do and what makes the pagination in the follow-ups necessary.
"Media?" Images and short video. Ask, because media introduces resumable uploads and a separate progress model, and if the answer is no you save fifteen minutes.
"Read receipts and typing indicators?" Ask specifically. Both are cheap-looking features with real costs: typing indicators produce a message every few keystrokes, which is a constant stream of tiny writes on the most expensive resource the phone has.
"End-to-end encryption?" Ask, and if the answer is yes, confirm whether it is in scope for this round. It changes the client design substantially — the follow-ups cover what — and if it is in scope you need to budget for it.
The scope to state back
In scope: send and receive text and media, one-to-one and small groups, message history on device, offline reading, offline sending that queues, correct ordering, and per-message delivery state.
Out of scope: voice and video calling, message search across the entire history, and the server's fan-out design.
The hard part, named early
"The interesting problem is that the connection drops constantly and the app is backgrounded often, so I need a design where a message the user sent is never lost and never sent twice, and where ordering does not depend on when things arrived."
Requirements and constraints
Functional requirements
- Send and receive messages in near real time while the app is open.
- Receive a notification when the app is backgrounded or killed.
- Read the full stored history with no network.
- Send with no network; the message queues and delivers later.
- Display messages in a stable, correct order on every device.
- Show per-message delivery state: sending, sent, delivered, read.
- Send and receive image and short video attachments.
The non-functional targets
| Target | Value | Note |
|---|---|---|
| Send-to-appear latency, both online | Under ~500 ms | Perceptible above roughly a second |
| Local echo latency | Under 16 ms | The sender's own message appears instantly |
| Conversation open to rendered | Under ~150 ms | Reading from the local database, not the network |
| Reconnect after a drop | Within ~5 s, backoff to a ceiling | Aggressive reconnect drains battery |
| Storage | Bounded, with a stated pruning rule | A heavy user accumulates gigabytes of media |
Note the second row. The sender's own message must appear at the moment they press send, before any network activity — the local echo. Waiting for a server acknowledgement makes a chat app feel broken on any connection worse than office Wi-Fi.
The four constraints
Network. This is the defining one. A mobile connection drops when the user walks into a lift, changes cell tower, moves from Wi-Fi to cellular, or the carrier's network address translation table expires an idle connection. A socket that has been silently dead for two minutes still looks open to your code until you write to it.
Battery. A persistent socket is not free. With a 30-second heartbeat the radio is woken 120 times an hour, and each wake keeps it in a high-power state for several seconds after the traffic ends. That can leave the radio energised a meaningful fraction of every minute.
Memory. A conversation with 50,000 messages must never be loaded into memory at once. Page it from the database, most recent first.
Storage. Text is small — 50,000 messages is a few megabytes. Media is not: a hundred received photos is 30–50 MB, and a heavy user reaches gigabytes. The pruning rule belongs in the design, not in a follow-up.
The constraint that shapes everything
The connection is intermittent and the app is often not running. Every design decision in the rest of this case study — the two-transport handoff, the persisted outbox, sequence-number ordering, gap detection — exists because of that one sentence.
Offline, explicitly: full history readable, sending queues with a "pending" indicator, typing indicators and read receipts suppressed, and a persistent banner. Nothing is lost, and nothing is silently dropped.
The transport decision
Chat needs sub-second delivery in both directions, and it needs to reach an app that is not running. No single transport does both. So the design uses two, and the handoff between them is the interesting part.
Why a socket, and why it cannot stay open
A WebSocket gives the lowest latency and the lowest per-message overhead, and it is bidirectional, which matters because the client sends as often as it receives. That makes it right for the foreground.
It cannot stay open in the background, for reasons that are not negotiable:
- When the app is backgrounded, the operating system suspends the process within seconds. Suspended code does not read from a socket.
- Under memory pressure the process is terminated outright.
- Even if you could keep it open, the battery cost above makes it indefensible for an app the user is not looking at.
Candidates who propose "keep the socket alive with a background service" are describing something the platform prevents, and saying so unprompted is one of the clearest platform judgement signals available in this round.
Push, and what a push actually is
A push notification travels over a single connection the operating system maintains for every app on the device. The cost is amortised across all of them, which is why it is roughly an order of magnitude cheaper than each app holding its own socket.
The trade is reliability. Delivery is best-effort, payloads are small, and the platform may throttle or coalesce a burst of notifications into one. Therefore: treat a push as a signal to fetch, not as the message itself. Carry enough in the payload to render the notification (sender, a snippet, the conversation identifier) and fetch the authoritative messages when the app next runs.
The handoff
The app moves between the two transports as it moves between foreground and background, and fetches whatever it missed on each return.
Detecting a dead socket, and reconnecting
A mobile connection often dies without a close frame, so the socket looks open. Two defences:
Application-level heartbeat. Send a ping every 30 seconds; if no pong arrives within about 10 seconds, treat the connection as dead and reconnect. Do not rely on the transport's own keep-alive, which can take many minutes to notice.
Reconnect with backoff and jitter. 1 s, 2 s, 4 s, 8 s, capped at 30–60 s, with a random offset. Without the cap, a long outage produces a reconnect storm; without the jitter, every client in a dropped cell reconnects in the same instant.
On reconnect, always resync: send the last known sequence number per conversation and let the server fill the gap. Never assume the socket picked up where it left off.
Offline: no socket, no push. Messages queue locally. On reconnect the outbox drains and the resync pulls what was missed, in that order.
Client architecture
The layered diagram from Step 3: the high-level client architecture again, with the pieces chat needs. The important property is that the network layer — both transports — never talks to the screen.
The layers
Presentation. A conversation-list screen and a conversation screen, each with a state holder observing a database query. The conversation screen observes messages where conversation_id = ? order by sort_key desc limit 50.
Data. A ChatRepository over: the device database (conversations, messages, outbox), file storage (attachments), a socket client, a push handler, and the HTTP client for history pagination and media.
The single source of truth
Every path writes to the same messages table and nothing else:
| Writer | What it writes |
|---|---|
| The user sending | A row with state pending and a client-generated identifier |
| The socket | Incoming messages, and acknowledgements that update state |
| The push handler | A placeholder or nothing, then triggers a fetch that writes real rows |
| History pagination | Older messages fetched on demand |
| The outbox drain | State transitions on rows the user sent |
Five writers, one table, one observed query. The screen has no idea which of the five caused the update it is rendering, and that is exactly the property that makes chat work offline.
The naive alternative keeps an in-memory list on the conversation screen and appends socket events to it. It works until the process is killed, the user rotates the device, or a push arrives while the screen is closed — at which point the list and the database disagree and messages appear to vanish and reappear.
Sending, stepped through
sendMessage(text): id = newClientId() // UUID, generated on device row = { id, conversation_id, text, author: me, state: PENDING, sort_key: localMax + 1 } // provisional ordering db.transaction { insert row into messages insert { id, attempt: 0 } into outbox } // returns here — the UI already updated from the DB changeThe function returns before any network call. The message is on screen in under 16 ms with a clock icon, and it is on disk, so killing the app now loses nothing.
Offline: every read path above is a local query, so the entire app except sending and receiving is unaffected.