Course Content
Mobile System Design Interview
11 sections · 23 lessons
Building blocks: persistence, caching, pagination and offline-first sync
Every mobile app stores things on the device. The design question is not whether but in which of the four stores, because putting data in the wrong one produces a specific, ugly failure months later.
This lesson follows data from the moment it lands on the device. It starts with where it is stored, then how it is cached and paged into a list, and ends with the single most important idea in the course — offline-first architecture — and the sync engine that keeps it honest.
The four stores
Key-value storage. Small named values: the selected theme, the last sync token, a feature-flag snapshot. Fast, trivial to use, loaded into memory whole. That last property is the trap — it is meant for kilobytes, not megabytes, and a key-value file that has grown to several megabytes slows app start because it is read before the first screen draws.
A relational database on device. The workhorse. Structured, queryable, indexable, and — the property that matters most for this course — observable: the user interface can subscribe to a query and be told when its results change. That is what makes the offline-first architecture later in this lesson possible.
File storage. Images, video, downloaded documents, anything measured in megabytes. Store the bytes in files and the metadata in the database, with the file path as a column. Never the reverse.
The secure store. A hardware-backed area for credentials: access tokens, refresh tokens, encryption keys. Nothing else belongs here — it is small and comparatively slow.
Choosing, in one table
| Data | Store | Why |
|---|---|---|
| User preferences, sync cursor | Key-value | Tiny, read at start-up |
| Feed posts, messages, file metadata | Database | Queryable, sortable, observable |
| Photos, videos, downloaded files | Files | Too large for a row; needs streaming |
| Auth tokens, encryption keys | Secure store | Needs hardware protection |
The migration problem
This is the part that bites, and it bites in production rather than in development.
Your app ships version 4 of the database schema. A user who has not opened the app in eight months launches it holding version 1. Your code must upgrade 1 → 2 → 3 → 4 on the main launch path, on a low-end device, with the user watching. If any step throws, the app crashes on launch, and — the constraint that defines this course — you cannot hotfix a shipped binary. A crash-on-launch migration bug means every affected user is stuck until they get an update from the store, which takes days.
Three rules that prevent it:
- Additive changes wherever possible. New nullable columns and new tables are safe. Renames, type changes, and dropped columns are the dangerous ones.
- Test every path, not the latest one. Keep a fixture database for each shipped schema version and run the migration chain against all of them in continuous integration.
- Have a fallback. If a migration fails, deleting and rebuilding the local cache is usually acceptable and always better than a crash loop. Data that cannot be rebuilt from the server — the outbox of unsent changes described below — is the exception, and needs its own recovery path.
Caching on the client
Some of what the stores hold is a cache. A cache on a phone is not a scaled-down server cache. A server cache trades memory for latency. A client cache trades storage for latency, battery, and the ability to work at all when there is no network.
Two layers, different jobs
The memory cache holds decoded, ready-to-use objects — a parsed model, a decoded bitmap. Access is instant. It dies with the process, which on a phone can be at any moment, so it is never the only copy of anything.
The disk cache holds bytes that survive restarts: response bodies, image files, the database rows themselves. Slower — tens of milliseconds rather than microseconds — but it is what makes a cold launch show content instead of a spinner.
The naive version keeps only a memory cache. It looks excellent in testing, where the app stays alive, and fails in the real world, where the user switches to another app for two minutes, the system reclaims your process, and they return to an empty screen and a full-network reload.
Sizing it for a device
A server cache is sized against a memory budget you control. A phone cache is sized against a budget shared with everyone.
- Memory cache: a common rule of thumb is roughly one-eighth of the app's available heap for images, and it must shrink on a low-memory warning. An unbounded memory cache is not a cache, it is a leak with a good reputation.
- Disk cache: pick a ceiling — 50–250 MB is a normal range for a media-heavy app — and evict against it. Also respect the operating system's cache directory semantics: content in a designated cache directory may be deleted by the system under storage pressure, which is correct behaviour and your code must survive it.
Eviction is least-recently-used in almost every case, bounded by total size rather than item count, because one 8 MB video thumbnail and one 4 KB JSON body should not count the same.
Invalidation: the three mechanisms
Time. Each entry carries a time-to-live. Feed pages might be 5 minutes; a currency rate 30 seconds; a user's avatar a week. Chosen per data type, never globally.
Validation. The server returns a version tag with a response. The client sends it back on the next request, and the server answers "unchanged" with an empty body instead of resending the payload. This still costs a round trip and a radio wake-up, but it saves the bytes — worthwhile for large responses, close to pointless for small ones.
Explicit. The client changed something, so it knows what is stale. After a successful edit, invalidate that item's entry directly rather than waiting for a timer.
Offline: the cache is the offline experience. This is why a cache policy that aggressively deletes on a schedule is a mistake — the entry you evicted at midnight is the one the user needed on the train at 08:00.
Pagination
Cached data usually reaches the screen as a list. A list on a phone can hold ten thousand items and a screen can show eight. Pagination is how those two facts are reconciled, and the choice between the two styles has a visible, user-facing consequence that backend candidates often miss.
Offset pagination, and how it breaks in public
Offset pagination asks for "20 items starting at 40". It is easy to implement and it is correct only while the list is not changing.
Watch it fail. A feed is sorted newest-first. The user loads page 1 (items 1–20). While they read, three new posts arrive at the top. The user scrolls, and the app asks for offset 20. But the list has shifted down by three, so items 18, 19 and 20 — which the user has already seen — are now at positions 21, 22 and 23, and come back again.
The user sees three duplicate posts. If items had been deleted instead, the same shift runs the other way and items are silently skipped, which is worse because nobody notices.
Cursor pagination
A cursor is an opaque token that means "the position just after this item". The client sends back exactly what the server gave it and asks for the next 20. Because the cursor names a position in the data rather than a count from the start, insertions and deletions above it do not shift anything.
GET /feed?limit=20→ { items: [...], next_cursor: "eyJ0IjoxNzA5..." }GET /feed?limit=20&cursor=eyJ0IjoxNzA5...→ { items: [...], next_cursor: "eyJ0IjoxNzA4..." }Two properties matter to the client and are worth asking for by name in an interview:
- The cursor is opaque. The client never parses it. That leaves the server free to change its encoding without shipping a new app — and shipping a new app takes days.
next_cursoris null at the end. An explicit end-of-list beats inferring it from a short page, which is ambiguous when the server returns fewer items for an unrelated reason.
Prefetch: hiding the latency
If you request the next page when the user reaches the last item, they wait. On a 4G network a page fetch is 150–400 ms, and a user flicking a list covers eight items in well under that.
Trigger the load when the user is within a threshold of the end — roughly one screen's worth, so five to ten items in a typical list. That converts a visible spinner into no perceptible pause. Set the threshold too high and you fetch pages nobody scrolls to, wasting the user's data and battery; two to three screens ahead is where it stops being worth it.
The placeholder
While a page is loading, show skeleton rows of the correct height rather than a spinner at the bottom. Correct height matters: rows that appear at zero height and then grow shift the content the user is reading, and that jump is one of the most complained-about defects in list-based apps.
Offline: paginate against local storage. The list reads from the database, so an offline user scrolls freely through everything already stored and hits a clear "you are offline" footer at the boundary instead of an error dialog.
Offline-first architecture
Every "offline" note so far has pointed at the same thing, and this is where it gets its name. It is the single most important idea in the course. If you take one thing from this lesson into an interview, take this: local storage is the source of truth, and the network is a synchroniser.
The naive architecture, and the moment it fails
Most first designs look like this: the screen asks the network for data, shows a spinner while it waits, and renders the response.
It works on office Wi-Fi. Then the user opens the app in a lift. There is no data to show, because the only copy lives on a server the phone cannot reach. So the screen shows an error — even though the app displayed this exact content ninety seconds ago and could have kept it.
The deeper problem is that the network is on the read path. Every screen is one dropped packet away from having nothing to draw.
Inverting it
Take the network off the read path entirely.
- The user interface observes a query against local storage.
- Local storage answers immediately, from whatever it has.
- Separately, a synchroniser fetches from the network and writes into local storage.
- The observed query fires again, and the screen updates.
The screen never calls the network and never knows whether one exists. That is the whole trick, and it is what makes offline behaviour fall out of the architecture rather than being bolted on per screen.
Writes: the outbox
Reads are half the story. When the user acts offline — sends a message, renames a file, likes a post — the change has to survive.
The outbox is a persisted queue of pending changes. A write does three things atomically: apply the change to local storage, mark that row as pending, and append an entry to the outbox. A background worker drains the outbox whenever there is a network, retrying with backoff and clearing the pending mark on success.
Because the outbox is on disk, it survives the app being killed. Because each entry carries a client-generated identifier, the server can recognise a retry and not duplicate the action.
Optimistic updates, and the rollback nobody writes
An optimistic update shows the result before the server confirms it. Tapping "like" fills the heart in under 16 ms rather than after a 400 ms round trip.
The half that gets skipped is failure. If the server ultimately rejects the change — the post was deleted, the quota is exceeded — you must revert the local row and tell the user something changed. Reverting silently is worse than never showing the optimistic state: the user saw it work.
The honest boundary: optimism is right for low-stakes, reversible actions. It is wrong for money, for orders, and for anything irreversible. The trading and hotel reservation case studies (Sections 6 and 8) are built on that line.
Sync and conflict resolution
The synchroniser in that loop needs a design of its own. Sync is the process of making local storage and the server agree. Conflict resolution is what happens when they disagree in a way that cannot be reconciled automatically — and the honest answer there is usually a product decision, not a technical one.
Full sync versus delta sync
Full sync downloads everything each time. It is trivially correct and gets expensive fast: a 5,000-message conversation is perhaps 2 MB of text, so syncing every launch spends the user's data allowance and their battery to learn that nothing changed.
Delta sync downloads only what changed since last time. The server hands the client an opaque change token with every response; the client stores it and sends it back on the next sync, and the server replies with the changes since that point plus a fresh token.
GET /changes?since=tok_8f2a1→ { changes: [ {op:"update", id:"f_31", ...}, {op:"delete", id:"f_88"} ], next_token: "tok_8f2c7", has_more: false }Three things this design has to get right:
- Deletions must appear as events. If a delete is only an absence, a delta sync can never learn about it. Servers usually keep tombstones for a retention window.
- The token must expire gracefully. If a client has been offline longer than the retention window, the server says "token too old" and the client falls back to a full sync. Handle it explicitly or you get a client that quietly stops receiving updates.
- Sync must be resumable.
has_moreplus a token lets a sync interrupted by process death continue where it stopped rather than restarting.
Detecting a conflict
A conflict is a local pending change to a record the server has also changed since your last sync. You detect it by comparing versions, not timestamps: the client keeps the version tag it last received, and the server rejects a write whose base version is stale.
Device clocks cannot do this job. Phone clocks drift, users set them manually, and time zones change under you — two devices can disagree by minutes. Version numbers from the server are monotonic and unambiguous; wall-clock time is neither.
The four resolutions
| Strategy | What happens | Fits |
|---|---|---|
| Last-write-wins | Latest server-stamped write survives | Preferences, read state, low-value data |
| Server-wins | Local change discarded | Data the server derives or authoritatively owns |
| Merge | Combine field by field, or use data types that merge by construction | Structured records with independent fields; collaborative text |
| Conflict copy or prompt | Keep both, or ask the user | Documents, anything the user authored |
Offline: the longer a device is offline, the more conflicts it comes back with. Design sync so a two-week-offline device is a normal case — full-sync fallback, batched outbox drain, and a conflict path that can handle more than one at a time.