Course Content
Mobile System Design Interview
11 sections · 23 lessons
Google Drive: requirements, the local data model and resumable transfers
Prompt: "Design a cloud file storage app — a Google Drive or Dropbox client."
This is the hardest sync problem in the course and the clearest demonstration of why offline-first is an architecture rather than a feature. Everything you have met so far — the outbox, delta sync, resumable transfer, conflict resolution — appears here at once, and they have to work together.
This lesson covers the scope, the constraints, the data model the whole design rests on, and how bytes move to and from the server. The next takes the sync engine, conflict resolution and the follow-ups.
The questions worth asking
"Browse, upload, download, or edit?" Browse, upload, and download. Push editing out of scope unless the interviewer insists — collaborative editing is a different problem with a different answer (operational transformation or conflict-free replicated data types), and it will consume the whole round.
"Which files are available offline — all of them, or ones the user selects?" This is the question that decides the design. A device has 64 GB of storage shared with everything else; a cloud account may hold two terabytes. Downloading everything is not possible, so files are available offline by selection, which immediately implies the metadata/content split in the local data model below.
"How large can a file be?" Ask for a number. It determines whether transfers must be chunked and resumable — and above a few tens of megabytes on a mobile connection, they must.
"Is sharing in scope?" Take it partially: shared folders appear in the tree and their permissions can change while you are syncing, which is a genuinely interesting client problem covered in the follow-ups. Push the sharing user interface out of scope.
"How many files in a typical account, and the largest realistic case?" A number here bounds your metadata store. Ten thousand files is comfortable; a million needs care about indexes and paging even for browsing.
The scope to state back
In scope: browsing the folder tree offline, marking files for offline availability, uploading in the background, downloading on demand, and two-way sync of changes.
Out of scope: document editing, real-time collaboration, the server's storage architecture, and search across file contents. Design Google Drive in System Design Interview designs the server side.
Requirements and constraints
Functional requirements
- Browse the full folder tree, offline, including folders whose contents are not downloaded.
- Mark a file or folder for offline availability, and unmark it.
- Open a downloaded file; download on demand for one that is not.
- Upload files from the device, continuing while the app is backgrounded.
- Sync changes both ways: remote changes appear locally, local changes reach the server.
- Rename, move, and delete, working offline and syncing later.
The non-functional targets
| Target | Value | Why |
|---|---|---|
| Folder open, offline | Under ~100 ms | It is a local indexed query |
| Metadata sync after reconnect | Seconds for a normal delta | Delta, not full |
| Upload resumability | Resumes at the last completed chunk | A 500 MB upload cannot restart |
| Offline file availability | 100% for pinned files | The user asked for these specifically |
| Storage footprint | Bounded and visible to the user | The app must never be the reason the phone is full |
The four constraints
Storage. The binding constraint. A 2 TB account against a device with perhaps 15 GB free is a ratio of over 100:1. Content is opt-in by definition, and the app needs an eviction policy and a user-visible storage screen.
Background execution limits. The second binding constraint, and the one that makes this hard. A 500 MB upload on a 5 Mbps connection takes roughly 13 minutes. Your app gets seconds of guaranteed execution after being backgrounded (Background work), so this transfer cannot be your code running in a loop. It must be handed to the platform's transfer service or scheduled as a deferred job — and even then it may be paused and resumed over hours.
Network. Uploads are long and mobile connections are short. Any transfer measured in minutes will be interrupted, so resumability is not an optimisation, it is a requirement.
Battery and data. Large transfers on cellular are expensive in both. A wifi-only default for large files, with an explicit override, is the defensible design.
Doing the metadata arithmetic out loud
A metadata row — identifier, parent, name, type, size, version, timestamps, flags — is roughly 200–300 bytes. So:
- 10,000 files ≈ 2–3 MB
- 100,000 files ≈ 20–30 MB
- 1,000,000 files ≈ 200–300 MB — large, but still feasible with care
Content for the same accounts could be hundreds of gigabytes. That two-to-three orders of magnitude gap is the justification for the entire architecture, and stating it with numbers is far more convincing than asserting that "metadata is small".
Offline, explicitly: browsing the whole tree works, pinned files open, unpinned files show a "not downloaded — connect to download" state rather than an error, uploads queue, and renames, moves and deletes apply locally and enter the outbox. The only thing genuinely unavailable is content that was never fetched.
Local data model
The arithmetic leads straight to the single modelling decision the entire design rests on: store the file tree separately from file content.
The naive model, and why it fails immediately
Model a file as one thing — a row that has a name, a parent, and its bytes. It is the obvious model and it fails at the first question: to show a folder, you would need the bytes of everything in it. Browsing a 2 TB account would mean downloading 2 TB.
Separating the two removes the problem entirely. Metadata is small enough to hold in full; content is fetched only when needed.
The schema
nodes( id text primary key, -- server id, stable parent_id text, -- null for the root; indexed name text, type enum FILE | FOLDER, size_bytes integer, version text, -- server-issued; drives conflict detection modified_at integer, -- SERVER time mime_type text, -- content state, per device -------------------------------- content_state enum NOT_DOWNLOADED | DOWNLOADING | LOCAL | STALE, local_path text, -- file on disk, or null pinned integer, -- user asked for this offline last_accessed integer, -- drives eviction -- local change state --------------------------------------- pending_op enum NONE | CREATE | RENAME | MOVE | DELETE | UPLOAD, local_version text)index on (parent_id, name)Content lives in files on disk. local_path points at it. The row is the truth about what exists; the file is the truth about what is downloaded.
Why content_state is a column and not an inference
It is tempting to infer state from whether local_path points at an existing file. Do not: checking the filesystem for every row makes a folder listing hundreds of file-system calls, and it cannot represent DOWNLOADING or STALE at all. Make it an explicit column, keep it correct in the same transaction as the file operation, and reconcile it against the filesystem occasionally rather than continuously.
STALE is the state that earns its place: the server has a newer version and the local copy is out of date. The user can still open the local copy offline — with a clear "older version" indicator — which is far better than refusing.
Eviction removes files, never nodes
When storage runs short, delete the file and set content_state = NOT_DOWNLOADED. Never delete the node. Deleting nodes leaves holes in the tree, so a folder appears to lose files that still exist in the account, which looks exactly like data loss to the user.
Evict by last accessed, never touching pinned files, and stop before the pinned set — if pinned content alone exceeds the budget, that is a message to the user, not a silent deletion of something they explicitly asked for.
Upload and download
With the tree on the device, the content has to move. Transfers here are minutes long on a connection that lasts seconds. Every decision follows from that.
Why a single request does not work
Uploading a 500 MB file as one request assumes an uninterrupted connection for the whole duration. On a phone, over roughly 13 minutes at 5 Mbps, that assumption fails often — a tunnel, a Wi-Fi handover, a screen lock. And when it fails at 480 MB, a single request has no way to resume: the whole transfer restarts, spending the user's data allowance twice.
Chunked, resumable transfers
Split the file into chunks — 4–8 MB is a common range, small enough that losing one is cheap and large enough that per-chunk overhead stays low.
1. POST /v1/uploads { name, parent_id, size, mime } → { upload_id, chunk_size }2. PUT /v1/uploads/{id}/chunks/{n} body = bytes for chunk n … repeat, in order or in parallel …3. POST /v1/uploads/{id}/complete { checksum } → { node_id, version } -- after any interruption -- GET /v1/uploads/{id} → { received_chunks: [0..57], expires_at }That last call is what makes resumption work: the client asks the server what it already holds and sends only what is missing. Persist upload_id and progress in the database, so resumption survives process death rather than only a network blip.
Upload sessions expire. Store expires_at, and if a resumed session has lapsed, start a fresh one rather than sending chunks into a session the server has forgotten.
Scheduling under background limits
A 13-minute transfer cannot be your own code. Three mechanisms, in order of preference:
- A system-managed transfer service. Hand the request to the operating system, which continues it after your process is suspended or killed and calls you back on completion. The correct choice for large files.
- A deferred job with constraints. "Requires unmetered network, prefer charging." Runs at the system's convenience, which may be hours. Correct for a bulk upload the user is not waiting on.
- Foreground work with a visible indicator. Only while the user is watching, and only for small files. It stops the moment they leave.
Progress that survives process death
Write progress to the database, not to a screen-owned object. Then the transfers screen renders from a query, a killed and relaunched app shows correct progress, and a notification can be updated from the same source. Coalesce the writes — update at most a few times a second rather than per chunk — which is the pattern from Throttling and coalescing updates applied to a different problem.
The wifi-only preference
Default large transfers to unmetered connections, with a per-transfer override and a clear queued state ("waiting for Wi-Fi") so the user is never left wondering why nothing is happening. Silence here reads as a bug; a labelled queue reads as a feature.
Offline: uploads queue with a visible pending state, downloads are unavailable and the file shows "not downloaded", and nothing fails silently.