Mobile System Design Interview

Course Content

Mobile System Design Interview

11 sections · 23 lessons

Google Drive: the sync engine, conflict resolution and follow-ups


The sync engine is the component the interviewer will pick for the deep dive. It has three parts — pulling remote changes, pushing local ones, and deciding what happens when both changed — and the ordering between them is where the difficulty lives.

This lesson builds the first two parts on top of the nodes table from the previous lesson, then gives the third its own treatment, because conflict resolution is where a reflex answer destroys somebody's work. It ends with the follow-ups this problem attracts.

Pulling: delta sync with a change token

Full sync of a hundred thousand nodes on every launch is 20–30 MB to learn that nothing changed. Instead the server issues an opaque change token:

Text
GET /v1/changes?token=chg_8812&limit=500→ { changes: [ { type: "upsert", node: {…} },                { type: "delete", id: "n_4471" },                { type: "move",   id: "n_9902", new_parent: "n_31" } ],    next_token: "chg_8830",    has_more: false }

Three requirements to state: deletions arrive as explicit events (an absence cannot be detected), has_more plus a token makes a long sync resumable across process death, and an expired token — a device offline longer than the server's retention — returns an explicit error that triggers a full resync rather than silent failure.

Apply a change set in one transaction so the tree is never observed half-updated.

Pushing: the outbox, with dependencies

Local changes go into an outbox (Offline-first architecture). The complication unique to a file tree is that operations depend on each other.

The user, offline, creates a folder "Invoices", moves three files into it, renames one, and deletes another. Draining that queue in arbitrary order fails: uploading a file into a folder the server has never heard of is a reference to a non-existent parent.

So the outbox is a dependency-ordered queue, not a simple one:

Text
outbox(seq, node_id, op, payload, depends_on_seq, state, attempts)

Drain in sequence order, and when an entry fails permanently, mark everything that depends on it blocked rather than continuing past it. Locally-created nodes get a temporary client identifier; when the server assigns a real one, rewrite the identifier across the tree and the remaining outbox entries in one transaction. That identifier remapping is a detail worth mentioning — it is where real implementations get subtle bugs.

Collapse redundant operations before sending: three renames of the same file offline is one rename, and a create followed by a delete is nothing at all.

DEVICELocal metadata DBwith a change tokenOutboxlocal mutationsConflict resolverSync enginepush then pullPOST /changessend the outboxGET /changes?since=tokenserver deltas409 conflictserver rejects a stale writeSERVERMetadata serviceauthoritativetoken12hand it to the resolverapply, advance the tokenpush before pull, always: sending localchanges first means the server's responsealready accounts for themthe resolver needs a policy —last-write-wins, or keep both as aconflicted copy
The change token is what makes sync incremental — without it every sync is a full download.

Where pull and push collide

The reconciliation rule, stated plainly: an incoming remote change to a node that has a pending local operation is a conflict. Everything else applies cleanly.

Order matters. Push the outbox first where possible, so the server's view includes your changes before you pull — that turns many potential conflicts into ordinary updates. Where a push cannot go first (the token is stale, or the outbox is blocked), pull into a staging area and reconcile explicitly rather than overwriting rows with pending local changes.

What triggers a sync

App foreground, a silent push hint from the server, a deferred job on a schedule, and an explicit pull-to-refresh. All four call the same idempotent routine, guarded so two triggers cannot run it concurrently.

Conflict resolution

Pushing first shrinks the problem; it does not remove it. Two devices edit the same file while both are offline. Both reconnect. There is no technical rule that makes this go away, and pretending otherwise is the mistake this part of the design exists to prevent.

Four resolutions, and what each costsPick a winner• Last write wins — silent data loss• Server wins — local edits vanish• Cheap, and occasionally wrongPreserve both• Keep both as a conflicted copy• Field-level merge where structured• Costly, but never destroys work
Prompting the user looks fair but arrives long after the edit, when they no longer remember either version.

Detecting it

Every node carries a server-issued version. The client stores the version it last saw and sends it with any change:

Text
PUT /v1/nodes/n_4471   If-Match: v_881   → 200  { version: "v_882" }          -- accepted   → 409  { current: { version: "v_889", … } }   -- someone else changed it

The 409 is the conflict. Note what is not used: timestamps. Device clocks disagree by minutes, so "whichever was modified later" cannot be computed reliably across devices. Server-issued versions are monotonic and unambiguous.

The four options, and what each costs

StrategyWhat the user experiencesWhere it fits
Last-write-winsOne version silently replaces the other; the loser is goneLow-value or derived data
Server-winsThe local change is discardedData the server owns
Conflict copyBoth survive: Report.docx and Report (conflict copy — Priya's iPhone, 30 Aug).docxUser-authored content
Prompt the userA dialog asking which to keepRare and important cases only

For file content, the conflict copy is the right default and is what mature products do. It never destroys work, it is understandable without explanation, and it leaves the resolution with the person who knows which version matters.

Metadata conflicts are different and often mergeable: a rename on one device and a move on another touch different fields and can both be applied. Say that — field-level merge where the fields are independent, conflict copy where they are not — because it shows the analysis rather than a reflex.

Prompting, and why it is usually wrong

Asking the user to choose seems respectful and is usually poor design. The conflict surfaces minutes or hours after the edit, out of context; the user does not remember which version had the paragraph they wanted; and a sync that returns forty conflicts produces forty dialogs. Prompt only when the resolution is genuinely irreversible and rare. Otherwise keep both and let the user resolve it when they are looking at the file.

Making conflicts less likely

Cheap measures worth naming: sync more often when the app is foregrounded, push local changes before pulling, and show a subtle indicator when a file the user is viewing has changed on the server. None of these eliminate conflicts, and a design that claims to have eliminated them is wrong — but they reduce the window meaningfully.

Offline: the longer offline, the more conflicts on return. Handle them in a batch — a single "3 files had conflicts, copies were created" summary — rather than one interruption each.

Follow-ups

Storage pressure and eviction

The app needs a budget, a policy, and a screen.

The eviction ladder under storage pressurePinnedoffline filesRecentlyopened contentThumbnailsand previewsCached,evicted firsttopbottomThe app needs a budget, a policy, and a user-visible screen showing all three.
Eviction runs from the bottom of the stack upward, so the user's explicit choices are the last thing to go.

Budget: pinned files (whatever they total) plus a cache ceiling for on-demand downloads — a few gigabytes, or a share of free space, whichever is smaller.

Policy: evict unpinned content by last accessed, delete the file, set content_state = NOT_DOWNLOADED, keep the node. If pinned content alone exceeds the budget, tell the user and let them unpin; never silently discard something they explicitly asked to keep offline.

Screen: a storage view showing what is used, by folder, with per-item and bulk clear actions. Also handle the operating system asking for space back: if the system reclaims your cache directory, the app must detect missing files on next access and re-download rather than show an error.

Very large files on a small device

A 4 GB video on a device with 3 GB free cannot be downloaded, and the app must say so before starting rather than failing at 90%. Check free space against file size plus a margin at the start of any download.

Two mitigations worth naming: stream media instead of downloading it whole, so a video can be played without ever fitting on disk; and offer a preview or thumbnail for formats where one exists. Also handle filesystem limits — some devices and formats cap individual file sizes, and that failure is confusing if unhandled.

Shared folders and permission changes mid-sync

The case that produces the strangest bugs. A folder is shared with the user, they pin it, and then access is revoked while a sync is running.

The design: treat permission as server-authoritative and arriving in the change feed. On a revocation, remove the subtree's nodes and delete downloaded content — leaving files they can no longer access is a real problem, not a convenience. Tell the user what happened ("Marketing is no longer shared with you") rather than letting a folder disappear silently. And handle in-flight operations: an upload into a folder you can no longer write to must fail cleanly with an explanation, and its outbox entry must be cancelled rather than retried forever.

Permission changes can also arrive during a sync, so the sync must tolerate a node becoming inaccessible partway through a change set — apply what is valid, drop what is not, and continue.

Encryption at rest

Both platforms encrypt device storage by default when the user has a passcode set, which covers a lost or stolen device. Beyond that, per-file encryption with keys held in the secure enclave protects against a compromised backup path or another app exploiting a vulnerability, at the cost of encrypt and decrypt time on every read and complexity around key rotation and multi-device access.

The honest position: rely on platform encryption plus correct file protection attributes for most products, and add application-level encryption when a compliance requirement or a threat model demands it — not by default. Security on device has the limits of client-side security.