Mobile System Design Interview

Course Content

Mobile System Design Interview

11 sections · 23 lessons

Building blocks: background work, media, security and safe rollout


Here is the constraint that most surprises engineers coming from the server side: when your app is not on screen, it is not running. Not "running slowly" or "running at low priority". Suspended, and eventually terminated.

This lesson covers the four building blocks that sit around the data layer: background work, images and media, security on the device, and the observability and rollout machinery that exists because a shipped binary cannot be hotfixed. Each one has a hard limit the operating system or the app store imposes, and each is a common interview probe.

What the operating system does when the user leavesForeground,runningBackground,seconds leftSuspended,no CPUTerminatedsilentlyNo callback fires on termination, so anything not persisted before suspension is gone.
Server intuitions fail here: off screen is not slow, it is stopped.

Background work

What actually happens when the user leaves

The user presses home. Your process is given a short window — on the order of tens of seconds — to finish what it is doing and save state. Then it is suspended: no threads run, no timers fire, no callbacks arrive. Later, under memory pressure, it is terminated outright, with no notification to your code.

That means a design that says "a background thread keeps the socket open and syncs every minute" is not a design. It describes something the operating system will not permit.

The three legitimate ways to run in the background

Deferred work schedulers. You hand the system a job description — "sync, needs network, prefer unmetered, prefer charging" — and the system runs it when convenient. Convenient may mean twenty minutes, or several hours if the device is idle in a pocket overnight. Both platforms batch these jobs across apps so the radio wakes once for many apps rather than many times for one.

Silent push. The server sends a data-only notification that wakes your app briefly. Useful and unreliable: delivery is best-effort, the platform throttles apps that overuse it, and a device in a deep idle state may hold the wake until it next comes alive.

System-managed transfers. Large uploads and downloads handed to an operating-system service that continues them after your process is gone, calling you back when finished.

Designing for it

Three rules follow.

  • Every background task is resumable. It can be killed at any point, so it must be able to restart from a persisted position rather than from the beginning.
  • Every background task is idempotent. It may run twice — once because the scheduler retried, once because you also triggered it on foreground.
  • Nothing the user is waiting on happens in the background. If it must be timely, do it in the foreground or push it to the user via a notification.

Offline: a scheduler constrained to "requires network" does not run at all while offline, which is the correct behaviour. The outbox from Offline-first architecture waits, and drains on the first scheduled run after connectivity returns.

Images and media

Background limits decide when work can run. Memory decides whether the app survives the work it does in the foreground, and images are the leading cause of out-of-memory crashes in real apps. The reason is a single arithmetic fact that people know and then fail to apply.

The image pipeline, and where the memory goesRequest by URLMemorycache hit?Disk cache hit?Fetch anddownsampleDraw atview sizeA 4000 by 3000 photo is 2 MB on disk and about 48 MB once decoded into pixels.
Decoding to the size of the view, not the size of the file, is what stops the out-of-memory crash.

The arithmetic

A JPEG on disk is compressed. A bitmap in memory is not. Decoded size is:

Text
width × height × 4 bytes

A 12-megapixel photo is 4032 × 3024, so decoded it occupies about 48 MB — regardless of the fact that the file was 3 MB. Hold five of those in a scrolling list and you have allocated 244 MB, which on most devices is past the point where the operating system kills your process.

Now the fix. That photo is being shown in a list row 400 × 300 points wide. At a 3× screen scale that is 1200 × 900 pixels, or about 4.3 MB decoded — eleven times less. Downsampling at decode time, rather than decoding fully and then scaling, is the difference between an app that crashes on a photo feed and one that does not.

The pipeline

Every serious image loader does the same six things, and naming them is what an interviewer is listening for:

  1. Resolve the request to a cache key, including the target size — the same URL at two sizes is two entries.
  2. Check the memory cache for a decoded bitmap.
  3. Check the disk cache for the encoded bytes.
  4. Fetch over the network if neither hits.
  5. Decode, downsampled to the target dimensions, off the main thread.
  6. Deliver to the view, if that view still wants it.

Step 6 is the one that gets skipped. In a recycling list, the row that requested the image may have been reused for a different item by the time the image arrives. Without a cancellation and identity check, the user sees the wrong picture flash into a row — the classic scrolling artefact — and you have spent battery decoding an image nobody sees. Scroll performance, in the news feed case study, draws this pipeline in full.

Progressive loading

An image arriving as a sudden pop is worse than one that fades in from something. The usual ladder: a solid placeholder of the right aspect ratio, then a tiny blurred thumbnail delivered inside the list payload itself, then the full image. The aspect ratio must come from the API — if the client does not know the dimensions before the bytes arrive, the row resizes when the image lands and the whole list jumps.

Video

Video is images with a deadline. Two rules carry most of the weight: keep one player instance and rebind it rather than creating one per row, and never decode more than one video at a time — hardware decoders are a limited resource and exhausting them fails in ways that look like unrelated bugs. Section 10 (YouTube App) goes into playback properly.

Offline: downloaded media lives in files with rows in the database pointing at them, so an offline user browses what has been fetched. Placeholders, not broken-image icons, for the rest.

Security on device

Client-side security has one honest framing: the device may belong to an attacker. Everything below raises the cost of an attack. None of it makes the client trustworthy.

What on-device security does and does not buyRaises the attacker's cost• Tokens in the keystore, not preferences• Short-lived access tokens• One refresh at a time, others wait• Pinning blocks casual interceptionDoes not make the client trusted• A rooted device reads your storage• Any check can be patched out• Pins expire and brick old builds• Authorisation belongs on the server
The honest framing is that the device may belong to the attacker; everything local only raises the price.

Storing credentials

Access tokens and refresh tokens go in the platform's secure store, which on modern devices is backed by a hardware security element and can require the user's biometric or passcode to unlock an item. They do not go in key-value preferences, which on a compromised device are a readable file.

Two properties worth choosing deliberately:

  • Availability. A token that is only readable after the user has unlocked the device once since boot is safer than one readable at any time — and it breaks background sync before first unlock. Pick knowingly.
  • Migration on restore. Decide whether credentials should survive a restore onto a different device. Usually they should not.

The token refresh race

Short-lived access tokens plus a long-lived refresh token is the standard design. The mobile failure mode is specific: a screen fires five requests at once, all five come back with "unauthorised", and all five independently call refresh. If the server rotates refresh tokens — issuing a new one and invalidating the old on every use — four of those five present a token that has been invalidated, and the user is logged out.

The fix is a single-flight refresh: the first failure takes a lock and refreshes, the other four wait on the same operation and retry with the resulting token. This is a small piece of code and a common interview probe, because it shows whether you have shipped an app.

Certificate pinning

Pinning means the app refuses connections whose server certificate does not match a key it was built with. It defeats an interception proxy and a rogue certificate authority.

It also creates a risk unique to mobile: the pin is compiled into a binary you cannot hotfix. If the server's certificate is rotated to a key the shipped app does not recognise, every installed copy stops working, and the fix is a store release plus however long users take to update. If you pin, pin to a certificate authority key rather than a leaf, ship at least one backup pin, and give the pin set an expiry after which the app falls back to normal validation rather than bricking.

Offline: an offline user cannot refresh an expired token. Decide what they see — reading cached content should keep working with a clear "signed-in session expired, reconnect to continue" state, rather than a hard logout that wipes local data the user may need.

Observability and rollout

The pinning problem is one instance of a wider rule. One constraint separates mobile from every other kind of engineering in this track, and it shapes the architecture rather than the tooling: you cannot hotfix a shipped binary.

Why you cannot hotfix a shipped binaryServer incident• Roll back in minutes• One version live at a time• Fix reaches everyone at onceMobile incident• Review plus adoption takes days• Many versions live for months• Old builds never update at all
Because the fix is slow, the architecture must carry the kill switch, the flag, and the staged rollout.

What that actually costs

A server bug is fixed and deployed in minutes. A mobile bug found an hour after release goes like this: build a fix, submit to the store, wait for review — hours to days — and then wait for users to update. Even with automatic updates, a meaningful fraction of your users will be on the broken version for a week, and a long tail runs versions many months old.

So every risky behaviour must be controllable without a release. That is not an operational nicety; it is a design requirement that changes what you build.

The four things that follow

Remote configuration. Values fetched at launch and cached: timeouts, page sizes, poll intervals, buffer targets. The cached copy must be usable offline and have a sane compiled-in default, because the first launch of a fresh install may have no network.

Kill switches. A boolean per risky feature. If the new checkout flow is failing, turn it off for everyone in a minute and fall back to the previous path, which is why the previous path has to stay in the binary for at least one release cycle.

Feature flags with targeting. The same mechanism, used to expose a feature to 1% of users, then 10%, then everyone — and to hold it at 10% while you read the numbers.

Staged rollout. Both app stores can release to a percentage of users and halt. A normal ladder is 1% → 5% → 20% → 50% → 100% over several days, with defined halt criteria: crash-free sessions dropping below the current release's baseline, a spike in a key error, or a fall in a core conversion metric.

What you must be able to measure

You cannot halt a rollout on a signal you do not collect.

SignalWhyA common target
Crash-free sessionsThe primary health metricOften 99.5%+
Application-not-responding / hang rateFreezes users notice but that do not crashTracked per release
Cold start timeFirst impression, and a ranking factor in storesUnder ~2 s
Frame drops on key screensPerceived qualityTracked per screen
Network error rate by endpointDistinguishes your bug from the server'sCompared release to release

Offline: analytics and crash reports must be written to a local queue and uploaded in batches when there is a network, never dropped. Batch them — one upload of two hundred events costs far less battery than two hundred uploads, for the radio-tail reason in The mobile constraints, quantified.