Course Content
Mobile System Design Interview
11 sections · 23 lessons
News Feed: requirements, API contract and client architecture
Prompt: "Design the news feed for a social app."
This is the first case study, and it establishes the template every later one follows. It is also the most transferable: an infinitely scrolling list of remote content with images is part of almost every app you will ever be asked to design.
Before reading on, spend the 45 minutes on it yourself. How to use this course has the protocol. This lesson covers the first half of the round — questions, requirements, the API and the architecture — and the next takes the deep dives and follow-ups.
The questions worth asking
"Which platforms, and do we support low-end devices?" Assume both platforms, and assume yes. That single answer sets the memory and rendering budget for everything below.
"Does the feed work offline?" Press for detail. The useful specification is: previously loaded posts are readable offline, reactions queue and send later, and new content requires a network. Say that back and get it confirmed.
"What media types?" Images certainly; ask about video explicitly, because autoplaying video changes the data, battery, and memory story completely.
"Is the feed chronological or ranked?" Ask once, then move on. Ranking is a server concern — the client receives an ordered list and renders it. Saying "that is the ranking service's job; from the client's side I need a stable order and a cursor" takes ten seconds and shows you know the boundary. Design a News Feed System in System Design Interview designs the server side if you want it.
"Are comments and reactions in scope?" Reactions yes, since they demonstrate the optimistic-update path. Comments as a separate screen, not inline, to keep the round focused.
The scope to state back
In scope: an infinite scrolling feed, pull-to-refresh, images and video, reacting to a post, opening a post detail screen, and offline reading of everything already fetched.
Out of scope: composing posts, the ranking algorithm, stories, direct messages, and the notification system.
Saying the out-of-scope list out loud is worth doing. It stops the interviewer wondering whether you forgot those things, and it protects your clock.
The assumption block
If the interviewer stays vague, commit:
"I'll assume both platforms, two OS versions back, low-end devices in scope, a server-ranked feed delivered with cursor pagination, images and video, and that previously fetched posts must be readable offline. Reactions queue offline; posting is out of scope."
Requirements and constraints
With scope agreed, write the requirements where both of you can see them. Functional first, then the four constraints that are present in every mobile design.
Functional requirements
- Scroll a feed of posts without a visible end.
- Pull down to refresh, without losing the reader's place.
- Render text, images, and video posts.
- React to a post, and see the reaction immediately.
- Read previously fetched posts with no network.
- Open a post from a deep link, arriving directly on its detail screen.
The non-functional targets
Numbers, because a requirement without one cannot be tested.
| Target | Value | Why this number |
|---|---|---|
| Scroll frame budget | 16.7 ms per frame at 60 Hz | One dropped frame is a visible stutter |
| Cold start to first content | Under ~2 s | Beyond that, users perceive the app as slow |
| Feed refresh latency | Under ~1 s on 4G | Roughly two round trips plus render |
| Disk cache ceiling | ~200 MB, evicted least-recently-used | Large enough for days of images, small enough not to be the reason the app is deleted |
| Memory | No allocation spike above the app's ceiling while scrolling | Exceeding it is process termination, not slowness |
On a 120 Hz display the frame budget halves to 8.3 ms. Design to 16.7 ms and note that the same techniques buy you the higher rate.
The four constraints
Network. Unreliable and variable: 20 ms on Wi-Fi, 150 ms on 4G, several seconds or nothing on a train. Reads must come from local storage so the feed renders regardless.
Battery. Each page fetch wakes the radio, and the radio stays awake for seconds after. Fetching 20 posts at a time rather than 5 means one wake instead of four for the same content. Video autoplay is the single largest battery item in a feed.
Storage. Images dominate. A feed image at 1080 × 1080 is roughly 150–400 KB as a compressed file; a few hundred of them is the 200 MB ceiling above.
Memory. The one that kills you. A decoded 1080 × 1080 bitmap is 1080 × 1080 × 4 ≈ 4.4 MB. Eight of those on screen is 35 MB, which is survivable. Eight full-resolution originals at 4032 × 3024 would be 390 MB, which is not. Images and media has the arithmetic.
The constraint that actually shapes the design
Of the four, smooth scrolling on a low-end device drives more decisions than anything else here. It is why images are downsampled at decode time, why display strings are computed when a post is written to the database rather than when a row is bound, why rows recycle, and why the next page is prefetched early. Name it as the governing constraint and the rest of your design has a spine.
Offline, explicitly: the feed reads from the device database, so an offline user scrolls everything previously fetched, sees a persistent "offline — showing saved posts" banner, can react (queued), and hits a clear end-of-cached-content footer instead of an error.
The API contract
Five minutes here, and one payload decision that matters more than all the others.
The feed endpoint
GET /v1/feed?limit=20&cursor=<opaque|absent>→ { "items": [ { "id": "p_8812", "server_time": "2026-08-30T09:14:02Z", "author": { "id": "u_41", "name": "Ravi Menon", "avatar_url": "…/u41_96.jpg" }, "text": "…", "media": [ { "type": "image", "url": "…/p8812_1080.jpg", "width": 1080, "height": 1350, "blur_placeholder": "…", "variants": { "480": "…", "1080": "…" } } ], "reactions": { "count": 214, "mine": null }, "comment_count": 31 } ], "next_cursor": "eyJ0IjoxNz…", "server_time": "2026-08-30T09:14:02Z" }Four things to point out as you write it: the cursor is opaque so the server can change its encoding without an app release; next_cursor is null at the end rather than inferred from a short page; the author and reaction counts are embedded so twenty posts is one request and not forty-one; and every response carries server time, because device clocks cannot be trusted to order anything.
The page-size trade-off
| Page size | Requests for 100 posts | Payload per request | Consequence |
|---|---|---|---|
| 5 | 20 | ~5 KB | 20 radio wakes; constant loading spinners |
| 20 | 5 | ~20 KB | Good balance; ~2–3 screens of content per page |
| 100 | 1 | ~100 KB | ~1 s+ to first render; most of it never scrolled to |
Twenty to twenty-five is the usual answer, chosen so that one page covers two or three screens — enough that the prefetch in Scroll performance always has time to complete.
The decision that matters most
Return image dimensions with every image URL.
Without them, the client does not know how tall a row is until the image has downloaded and decoded. So the row is laid out at some guessed height, the image arrives 400 ms later, the row resizes, and everything below it jumps — while the user is reading it. This is the single most complained-about defect in feed apps, and it is fixed entirely in the API.
With width and height the client computes the exact row height from the aspect ratio before any bytes arrive, reserves the space, and drops the image in without moving anything.
The blur_placeholder — a very small encoded thumbnail, tens of bytes, carried inline — completes it: the reserved space shows a blurred approximation immediately rather than a grey rectangle. And variants lets a low-end device or a cellular connection fetch the 480 px version instead of the 1080 px one, which is roughly a fifth of the bytes.
Offline: every field above is stored, so an offline render is byte-identical to an online one from cache. Reactions post to POST /v1/posts/{id}/reactions with a client-generated idempotency key so a queued retry cannot double-count.
Client architecture
With the payload fixed, draw the app that consumes it. This is the layered diagram from Step 3: the high-level client architecture with feed-shaped labels. Draw the same shape every time; change what is in the boxes.
The layers, filled in
Presentation. A feed screen and a state holder. The state holder exposes one stream: a list of posts, plus a loading state for the page being fetched. It calls two methods on the repository — observeFeed() and loadMore() — and knows nothing else.
Data. A FeedRepository, a local store (the device database plus the image file cache), and a remote source (the HTTP client from The network layer).
There is no meaningful domain layer here. Say that rather than inventing one: "there are no business rules in a feed beyond ordering, so I'd keep the domain layer out and add it if reaction rules get complicated."
Why the feed renders from the database even when online
The naive design has the repository return the network response directly, and write to the database on the side as a cache. It behaves correctly until a second source of truth appears — a push notification updating a reaction count, or the post detail screen editing the same post — and then the feed row and the detail screen disagree, because each holds its own copy.
Reading everything through the database removes the possibility. Any writer — the feed fetch, the detail screen, a push handler, the outbox — writes to one table, and every observer sees the same value within a frame or two. The cost is a few milliseconds of database read per update, which is far below the 16.7 ms frame budget.
The load-more path, stepped through
state holder: loadMore() → repository.loadMore() → if a load is already in flight: return (dedupe) → cursor = db.readCursor(feedId) → response = api.feed(limit=20, cursor=cursor) → db.transaction { upsert posts // content, keyed by post id append feed_entries // ordered positions → post id write next_cursor } → db emits a change → observeFeed() emits → screen rendersThe repository returns nothing useful. Its job is to write into the database; the update reaches the screen through the observation, not through the return value. That inversion is the architecture.
Two tables, not one
Content and ordering are stored separately:
posts(id, author_id, text, media_json, reaction_count, my_reaction, updated_at)feed_entries(feed_id, position, post_id, page_cursor)
The same post can appear in several feeds without being duplicated, a re-rank changes only feed_entries, and a reaction updates posts once and every feed containing it updates together.