Mobile System Design Interview

Course Content

Mobile System Design Interview

11 sections · 23 lessons

Google Drive: requirements, the local data model and resumable transfers


Prompt: "Design a cloud file storage app — a Google Drive or Dropbox client."

This is the hardest sync problem in the course and the clearest demonstration of why offline-first is an architecture rather than a feature. Everything you have met so far — the outbox, delta sync, resumable transfer, conflict resolution — appears here at once, and they have to work together.

This lesson covers the scope, the constraints, the data model the whole design rests on, and how bytes move to and from the server. The next takes the sync engine, conflict resolution and the follow-ups.

Scoping a file storage clientDrive scopeBrowse, or edit too?Offline access opt-in?How large are files?Shared folders?Which conflict policy?
Whether offline editing is in scope decides whether conflict resolution is the centre of this design.

The questions worth asking

"Browse, upload, download, or edit?" Browse, upload, and download. Push editing out of scope unless the interviewer insists — collaborative editing is a different problem with a different answer (operational transformation or conflict-free replicated data types), and it will consume the whole round.

"Which files are available offline — all of them, or ones the user selects?" This is the question that decides the design. A device has 64 GB of storage shared with everything else; a cloud account may hold two terabytes. Downloading everything is not possible, so files are available offline by selection, which immediately implies the metadata/content split in the local data model below.

"How large can a file be?" Ask for a number. It determines whether transfers must be chunked and resumable — and above a few tens of megabytes on a mobile connection, they must.

"Is sharing in scope?" Take it partially: shared folders appear in the tree and their permissions can change while you are syncing, which is a genuinely interesting client problem covered in the follow-ups. Push the sharing user interface out of scope.

"How many files in a typical account, and the largest realistic case?" A number here bounds your metadata store. Ten thousand files is comfortable; a million needs care about indexes and paging even for browsing.

The scope to state back

In scope: browsing the folder tree offline, marking files for offline availability, uploading in the background, downloading on demand, and two-way sync of changes.

Out of scope: document editing, real-time collaboration, the server's storage architecture, and search across file contents. Design Google Drive in System Design Interview designs the server side.

Requirements and constraints

Functional requirements

  1. Browse the full folder tree, offline, including folders whose contents are not downloaded.
  2. Mark a file or folder for offline availability, and unmark it.
  3. Open a downloaded file; download on demand for one that is not.
  4. Upload files from the device, continuing while the app is backgrounded.
  5. Sync changes both ways: remote changes appear locally, local changes reach the server.
  6. Rename, move, and delete, working offline and syncing later.

The non-functional targets

TargetValueWhy
Folder open, offlineUnder ~100 msIt is a local indexed query
Metadata sync after reconnectSeconds for a normal deltaDelta, not full
Upload resumabilityResumes at the last completed chunkA 500 MB upload cannot restart
Offline file availability100% for pinned filesThe user asked for these specifically
Storage footprintBounded and visible to the userThe app must never be the reason the phone is full

The four constraints

Storage. The binding constraint. A 2 TB account against a device with perhaps 15 GB free is a ratio of over 100:1. Content is opt-in by definition, and the app needs an eviction policy and a user-visible storage screen.

Metadata fits on the device; content does notMetadata, always synced• Names, sizes, ids, parents• Roughly a kilobyte per file• A full tree costs megabytesContent, opt-in only• A 2 TB account, 15 GB free• A ratio over one hundred to one• Needs eviction and a storage screen
Doing this arithmetic out loud is what justifies syncing the whole tree while downloading almost none of it.

Background execution limits. The second binding constraint, and the one that makes this hard. A 500 MB upload on a 5 Mbps connection takes roughly 13 minutes. Your app gets seconds of guaranteed execution after being backgrounded (Background work), so this transfer cannot be your code running in a loop. It must be handed to the platform's transfer service or scheduled as a deferred job — and even then it may be paused and resumed over hours.

Network. Uploads are long and mobile connections are short. Any transfer measured in minutes will be interrupted, so resumability is not an optimisation, it is a requirement.

Battery and data. Large transfers on cellular are expensive in both. A wifi-only default for large files, with an explicit override, is the defensible design.

Doing the metadata arithmetic out loud

A metadata row — identifier, parent, name, type, size, version, timestamps, flags — is roughly 200–300 bytes. So:

  • 10,000 files ≈ 2–3 MB
  • 100,000 files ≈ 20–30 MB
  • 1,000,000 files ≈ 200–300 MB — large, but still feasible with care

Content for the same accounts could be hundreds of gigabytes. That two-to-three orders of magnitude gap is the justification for the entire architecture, and stating it with numbers is far more convincing than asserting that "metadata is small".

Offline, explicitly: browsing the whole tree works, pinned files open, unpinned files show a "not downloaded — connect to download" state rather than an error, uploads queue, and renames, moves and deletes apply locally and enter the outbox. The only thing genuinely unavailable is content that was never fetched.

Local data model

The arithmetic leads straight to the single modelling decision the entire design rests on: store the file tree separately from file content.

The naive model, and why it fails immediately

Model a file as one thing — a row that has a name, a parent, and its bytes. It is the obvious model and it fails at the first question: to show a folder, you would need the bytes of everything in it. Browsing a 2 TB account would mean downloading 2 TB.

Separating the two removes the problem entirely. Metadata is small enough to hold in full; content is fetched only when needed.

The schema

Text
nodes(    id              text primary key,     -- server id, stable    parent_id       text,                 -- null for the root; indexed    name            text,    type            enum FILE | FOLDER,    size_bytes      integer,    version         text,                 -- server-issued; drives conflict detection    modified_at     integer,              -- SERVER time    mime_type       text,    -- content state, per device --------------------------------    content_state   enum NOT_DOWNLOADED | DOWNLOADING | LOCAL | STALE,    local_path      text,                 -- file on disk, or null    pinned          integer,              -- user asked for this offline    last_accessed   integer,              -- drives eviction    -- local change state ---------------------------------------    pending_op      enum NONE | CREATE | RENAME | MOVE | DELETE | UPLOAD,    local_version   text)index on (parent_id, name)

Content lives in files on disk. local_path points at it. The row is the truth about what exists; the file is the truth about what is downloaded.

Metadata tree — small, always synced/Work/Photos/report.pdftrip.jpga few KB for a whole account — sync it eagerly and the file browser worksofflineContent — large, fetched on demandRemote blobsLocal cacheLRU, size-cappeddownload when openedwhy the split matters on mobileThe user has 60 GB in the cloud and 4 GB free on the phone. Syncing metadatagives them a complete, browsable, searchable file list; syncing content would fillthe device on the first launch."Available offline" then becomes a per-file user choice, pinned in the metadata and honoured by the cache eviction policy.
Metadata is kilobytes and content is gigabytes — syncing them on the same schedule is what fills a user's phone.

Why content_state is a column and not an inference

It is tempting to infer state from whether local_path points at an existing file. Do not: checking the filesystem for every row makes a folder listing hundreds of file-system calls, and it cannot represent DOWNLOADING or STALE at all. Make it an explicit column, keep it correct in the same transaction as the file operation, and reconcile it against the filesystem occasionally rather than continuously.

STALE is the state that earns its place: the server has a newer version and the local copy is out of date. The user can still open the local copy offline — with a clear "older version" indicator — which is far better than refusing.

Eviction removes files, never nodes

When storage runs short, delete the file and set content_state = NOT_DOWNLOADED. Never delete the node. Deleting nodes leaves holes in the tree, so a folder appears to lose files that still exist in the account, which looks exactly like data loss to the user.

Evict by last accessed, never touching pinned files, and stop before the pinned set — if pinned content alone exceeds the budget, that is a message to the user, not a silent deletion of something they explicitly asked for.

Upload and download

With the tree on the device, the content has to move. Transfers here are minutes long on a connection that lasts seconds. Every decision follows from that.

A chunked, resumable transferOpenupload sessionSplitinto chunksSend andrecord offsetResume fromlast offsetFinalisethe fileProgress lives in the database, so process death resumes rather than restarts.
Transfers last minutes on connections that last seconds, so the unit of work has to be smaller than the file.

Why a single request does not work

Uploading a 500 MB file as one request assumes an uninterrupted connection for the whole duration. On a phone, over roughly 13 minutes at 5 Mbps, that assumption fails often — a tunnel, a Wi-Fi handover, a screen lock. And when it fails at 480 MB, a single request has no way to resume: the whole transfer restarts, spending the user's data allowance twice.

Chunked, resumable transfers

Split the file into chunks — 4–8 MB is a common range, small enough that losing one is cheap and large enough that per-chunk overhead stays low.

Text
1. POST /v1/uploads               { name, parent_id, size, mime }                                  → { upload_id, chunk_size }2. PUT  /v1/uploads/{id}/chunks/{n}   body = bytes for chunk n   … repeat, in order or in parallel …3. POST /v1/uploads/{id}/complete { checksum }                                  → { node_id, version }   -- after any interruption --   GET  /v1/uploads/{id}          → { received_chunks: [0..57], expires_at }

That last call is what makes resumption work: the client asks the server what it already holds and sends only what is missing. Persist upload_id and progress in the database, so resumption survives process death rather than only a network blip.

Upload sessions expire. Store expires_at, and if a resumed session has lapsed, start a fresh one rather than sending chunks into a session the server has forgotten.

Scheduling under background limits

A 13-minute transfer cannot be your own code. Three mechanisms, in order of preference:

  1. A system-managed transfer service. Hand the request to the operating system, which continues it after your process is suspended or killed and calls you back on completion. The correct choice for large files.
  2. A deferred job with constraints. "Requires unmetered network, prefer charging." Runs at the system's convenience, which may be hours. Correct for a bulk upload the user is not waiting on.
  3. Foreground work with a visible indicator. Only while the user is watching, and only for small files. It stops the moment they leave.

Progress that survives process death

Write progress to the database, not to a screen-owned object. Then the transfers screen renders from a query, a killed and relaunched app shows correct progress, and a notification can be updated from the same source. Coalesce the writes — update at most a few times a second rather than per chunk — which is the pattern from Throttling and coalescing updates applied to a different problem.

The wifi-only preference

Default large transfers to unmetered connections, with a per-transfer override and a clear queued state ("waiting for Wi-Fi") so the user is never left wondering why nothing is happening. Silence here reads as a bug; a labelled queue reads as a feature.

Offline: uploads queue with a visible pending state, downloads are unavailable and the file shows "not downloaded", and nothing fails silently.