Course Content
Mobile System Design Interview
11 sections · 23 lessons
Building blocks: the mobile constraints and the network
A mobile app runs on a device that is short of power, short of memory, short of network, and can be killed by the operating system at any moment. Every design decision in the rest of this course is a trade against one of those four.
This lesson puts numbers on those constraints, then builds the component that meets the worst of them head on: the network layer. It ends with the decision an interviewer always pushes on — how the client learns that something changed on the server.
Constraint 1: the network is not a wire
On a good office Wi-Fi network a round trip to a nearby server takes roughly 20–50 ms. On a loaded 4G cell it is more often 50–150 ms. On a congested train, in a lift, or at the edge of coverage it can be 500–2,000 ms, or the request never completes at all.
The number that surprises backend engineers is the cost of the first request. Opening a fresh connection needs a name lookup, a transport handshake, and an encryption handshake — about three round trips before a single byte of your payload moves. At 150 ms per round trip that is roughly 450 ms of waiting before the request has even been sent.
Constraint 2: battery is a budget you spend
A typical phone battery holds somewhere around 11–19 watt-hours. The cellular radio, the screen, and the GPS receiver are the three large consumers, and the radio has a property that catches people out: after you send anything, it stays in a high-power state for several seconds waiting for more traffic. That trailing period — usually called the tail — means a 1 KB request can cost nearly as much energy as a 100 KB one.
So the rule is not "send less data". It is send less often, and batch what you send.
Constraint 3: memory, and the process death that follows
Each app gets a memory allowance far below the device's total. Exceed it and the operating system terminates your process without warning and without running your cleanup code. The single most common cause is images: a decoded bitmap costs width × height × 4 bytes, so one 12-megapixel photo occupies roughly 48 MB in memory regardless of how small the JPEG file was on disk.
Constraints 4 and 5: storage and heat
Device storage is shared with photos, music, and every other app. A cache that grows without a ceiling will eventually be the reason a user deletes your app. And sustained CPU or GPU load makes the device hot, at which point the operating system throttles the processor — your app gets slower precisely when it is working hardest.
Offline, in one line: none of these constraints disappear offline; the network one becomes absolute, and everything the app can still do has to come from local storage.
The network layer
The first constraint lands on one component. The network layer is the one component in the app that turns "I need this data" into bytes on the wire and back again. Everything else in the client depends on it, which is why a sloppy one produces symptoms that look like bugs everywhere else.
The naive version, and how it fails
The naive network layer creates a fresh client for each request and calls it. Three things go wrong.
It pays the handshake every time. Name lookup, transport handshake, encryption handshake — about three round trips. On a 150 ms link that is ~450 ms of dead time added to every call. A screen that makes four sequential calls now takes nearly two seconds before any of its own work begins.
It duplicates work. A user taps a row twice, or two screens both need the profile. Two identical requests go out, two responses come back, and whichever finishes last wins — which may be the older one.
It never gives up, or gives up instantly. With no timeout, a request on a dead connection hangs until the operating system decides otherwise, which can be minutes. With an aggressive timeout, a slow-but-working network looks broken.
What a real network layer contains
Connection reuse. One shared client, one connection pool. Subsequent requests to the same host skip the handshake entirely, turning that 450 ms into roughly one round trip.
Three separate timeouts. A connect timeout (~10 s), a read timeout for the gap between bytes (~15–30 s), and a total deadline for the whole call. Only the third protects a user from a server that dribbles a byte every 20 seconds forever.
Retry with exponential backoff and jitter. Retry only what is safe to repeat. Wait 1 s, then 2 s, then 4 s, and add a random offset to each so that ten thousand phones coming back onto a network do not all retry in the same millisecond.
Request deduplication. Keep a map of in-flight requests keyed by method plus path plus parameters. A second identical request attaches to the first rather than issuing its own.
Cancellation. When the user navigates away, the work for the screen they left must stop. Without this, a fast scroll through ten screens leaves ten live requests competing for the same radio and the same thread pool.
Distinguishing "offline" from "broken"
A network layer that treats no-connectivity as a server error will retry into an aeroplane mode that has been on for an hour, burning battery for nothing. Watch the connectivity state instead: when there is no route to the network, stop retrying, mark the request pending, and flush the queue on the reconnect callback.
Getting data from the server: five options
A network layer sends requests the client decides to make. The harder question is how the client learns that something changed on the server without asking. There are five ways, and choosing between them is one of the two or three decisions an interviewer will always push on. The mobile answer differs from the backend answer for one reason: four of the five stop working the moment the app is backgrounded.
The five
Polling. The client asks on a timer. Simple, works everywhere, and wasteful: at a 30-second interval, the average staleness is 15 seconds and the app wakes the radio 120 times an hour, mostly to be told nothing changed.
Long polling. The client asks and the server holds the request open until it has something or a timeout fires. Latency drops close to real time without a new protocol, but each hanging request holds a connection, and mobile networks routinely kill idle connections after 30–60 seconds, forcing a reconnect cycle anyway.
Server-sent events. One long-lived HTTP response over which the server pushes a stream of text events. Server-to-client only, reconnects automatically with a last-event ID, and is much simpler than a full socket. A good fit when the client never needs to push.
WebSockets. A full two-way connection. Lowest latency, lowest per-message overhead, and the right answer for chat and for streaming prices. The costs are real: you own reconnection and resubscription, you need an application-level heartbeat because a dead mobile connection often looks alive, and the socket dies when the app is backgrounded.
Push notifications. A message delivered by the platform's own always-on channel. This is the only option in the list that works when your app is not running. It is also the least reliable: delivery is best-effort, payloads are small (on the order of a few kilobytes), and the platform may throttle or coalesce a burst.
The comparison that decides it
| Latency | Battery | Works backgrounded | Client complexity | |
|---|---|---|---|---|
| Polling | interval ÷ 2 | Poor at short intervals | No | Very low |
| Long polling | Near real time | Medium | No | Low |
| Server-sent events | Near real time | Medium | No | Low |
| WebSocket | Lowest | Poor if chatty; good if quiet | No | High |
| Push notification | Seconds, not guaranteed | Excellent — the platform pays | Yes | Medium |
The mobile-specific move
Because only push survives backgrounding, real apps use two transports and hand off between them: a socket or stream while the user is looking at the screen, push while they are not, and a fetch-on-wake that pulls whatever was missed. Section 5 (Chat App) works this through for chat; Section 6 (Stock Trading App) does it for market data, where the answer is different because a stale price is dangerous rather than merely stale.
Offline: all five fail identically with no network. What differs is recovery. Polling recovers by itself on the next tick. A socket needs an explicit reconnect with backoff and a resubscribe. Push may deliver a notification that arrives before your data does, so the notification handler must be able to fetch, not just display.