System Design Interview

Course Content

System Design Interview

31 sections · 71 lessons

YouTube: scope, scale and the upload path


The prompt: "Design YouTube."

Attempt it for 45 minutes before reading. Draw the upload path and the playback path separately — they share almost nothing, and candidates who try to draw one diagram for both get stuck within five minutes.

This lesson scopes the problem, runs the estimate that turns the playback design from a choice into a consequence, and then builds the first of the two systems: the upload path, where the engineering is. The playback path, where the money is, follows in YouTube: streaming, cost optimisation and failure handling.

Two systems sharing one promptThe upload path• Minutes to hours of processing• Transcode into many renditions• Failure can be retried invisiblyThe playback path• Milliseconds, and it must not fail• Served almost entirely by the CDN• This is where the money goes
The expensive work happens once, offline; the cheap work happens a billion times, and that shapes everything.

Why video is a different shape of problem

Every system so far has moved records: a post, a message, a notification. Records are kilobytes. Video is gigabytes, and that single change of unit rewrites the design.

Two consequences follow immediately, and stating them early frames the whole conversation:

  1. The bytes never go through your application servers. A 600 MB upload streamed through a web tier would occupy a process for minutes. Uploads go directly to blob storage; playback comes directly from a content delivery network.
  2. Cost is a first-class design constraint. In every earlier case study, cost was a footnote. Here, egress bandwidth is the dominant line item, and the estimate below shows the arithmetic that makes it so.

The five questions

1. Upload and playback only, or recommendations and search too? Recommendations are a machine-learning system (Machine Learning System Design Interview covers that shape) and search is another. Scope to upload, processing, and playback, and say why: those three contain the hard parts specific to video.

2. Which resolutions and devices? This decides the encoding ladder — the set of quality levels every uploaded video is converted into. Six levels multiply storage and transcoding cost by six, and the estimate below computes that directly.

3. Is live streaming in scope? Say no for the main design, then discuss it as a follow-up. Live has a latency budget measured in seconds and cannot use the batch upload pipeline described below at all. Conflating the two is the most common way to lose control of this question.

4. Global audience? If yes, the content delivery network is not an optimisation, it is the serving architecture. Assume yes.

5. What is the read-to-write ratio, in bytes? Push for this one. Video platforms are overwhelmingly read-heavy — the estimate below computes roughly 150 bytes served for every byte ingested — and that ratio is the justification for spending nearly all of the engineering effort on the delivery path.

Requirements and scale

The estimate is the argument for the architecture. Do it properly and the CDN-centric design in the streaming path stops being a choice and becomes a consequence.

The estimate that forces a CDN5 M uploadersGlobal ingest500 hours a minuteQueue, not syncabout 1 PB a dayTiered storage150 Tbps at peakOrigincannot serve5x real time eachParallel chunksNumberConsequenceDaily usersUploadsStorageEgressTranscode
At this egress the CDN stops being an optimisation and becomes the only affordable way to serve the bytes.

Requirements

Functional: upload a video with metadata; process it into multiple resolutions; play it back on any device with quality that adapts to the network; resume an interrupted upload; delete a video.

Non-functional: playback starts in under 2 seconds; playback does not stall; uploads are resumable and never silently lost; processing completes within minutes for typical videos; and cost per hour delivered stays within budget, which is a real requirement here.

The numbers

Invented figures for a large video platform. A day is 100,000 seconds (see Rounding aggressively and staying fast).

QuantityAssumption
Uploads per day1 million
Average video length10 minutes
Source bitrate8 Mbps (1080p)
Daily active viewers200 million
Watch time per viewer per day30 minutes
Average delivered bitrate2 Mbps (mobile-heavy mix)

Ingest. A 10-minute video at 8 Mbps is 8 × 600 = 4,800 megabits = 600 MB.

1M uploads × 600 MB = 600 TB/day of raw ingest = 6 GB/s average

Stored output. An encoding ladder of six renditions — roughly 0.1, 0.3, 0.7, 1.5, 3, and 6 Mbps — sums to about 11.6 Mbps, so 10 minutes of output across all renditions is 11.6 × 600 = 6,960 megabits ≈ 870 MB, call it 1 GB with audio tracks and thumbnails.

1M × 1 GB = 1 PB/day stored = 365 PB/year, growing forever

Egress — the number that explains the design.

200M viewers × 0.5 hours = 100 million watch-hours/day

2 Mbps × 3,600 s = 7,200 megabits/hour = 900 MB per watch-hour

100M × 900 MB = 90 PB/day delivered

90 PB ÷ 100,000 s = 900 GB/s average ≈ 7 Tbps, and roughly 15 Tbps at a 2× peak

The ratio: 90 PB out against 0.6 PB in is 150 bytes served for every byte ingested.

Transcoding compute

Converting one 10-minute video into six renditions takes, very roughly, a few minutes of CPU per rendition depending on codec and quality preset — say 18 core-minutes per video in total.

1M videos × 18 core-minutes = 18 million core-minutes/day = 300,000 core-hours/day

300,000 ÷ 24 = about 12,500 cores running continuously

Encoding cost varies by an order of magnitude with codec choice and preset, so present this as a sizing exercise rather than a measurement. The useful conclusion is that transcoding is a large, elastic, batch workload — a fleet, not a service.

The upload path

This is the first deep dive, and the first half of the one new idea this section owns: modelling video processing as a directed acyclic graph of tasks, paired with the CDN economics in the streaming path.

Getting the bytes in

The naive version — one HTTP POST of 600 MB to an application server, which forwards it to storage — fails three ways. The connection drops on a mobile network and the whole upload restarts. The application server holds a process for the duration. And the server's bandwidth becomes the bottleneck at 6 GB/s of aggregate ingest.

The replacement has two parts:

Chunked, resumable upload. The client splits the file into chunks of 5 to 10 MB and uploads them individually, recording which succeeded. A dropped connection costs one chunk, not 600 MB. On resume the client asks the server which chunks it already has and sends the rest.

Pre-signed upload URLs. The application server issues a short-lived, cryptographically signed URL that authorises writing one specific object to blob storage. The client uploads directly to storage; the bytes never touch the application tier. The server's job shrinks to issuing URLs and recording completion — a few hundred bytes of work per 600 MB moved.

Processing as a directed acyclic graph

Once the upload completes, storage emits an event and the processing pipeline starts. The work is not one job; it is a set of tasks with dependencies.

The stages, in order:

  1. Inspect — probe the container and codec, read duration and resolution, reject anything malformed. Cheap, and it fails fast before expensive work starts.
  2. Split — cut the source into segments of a few seconds each, aligned to group-of-pictures boundaries (the points where a video stream contains a complete frame that later frames are described relative to). Splitting anywhere else produces segments that cannot be decoded independently.
  3. Encode — the wide part of the graph. Every segment × every rendition is an independent task. A 10-minute video split into 100 six-second segments across 6 renditions is 600 independent encode tasks.
  4. Side branches — thumbnail extraction, audio track encoding, watermarking, subtitle generation, and content-safety scanning all hang off the inspect stage and run in parallel with encoding.
  5. Merge and package — reassemble each rendition's segments, write the streaming manifest (the index file listing every rendition and every segment), and produce the final playable output.
  6. Publish — write metadata to the database, mark the video available, warm the content delivery network for popular uploaders.

An orchestrator holds the graph state; tasks are queued to a distributed log and consumed by a worker fleet that scales with the backlog.

Uploadpresigned URLRaw storeInspectcodec, length1080p720p480paudiothumbnailsPackagerHLS / DASH manifestCDNThe five branches are independent, so they run in parallel and a failure in one resolution does not block the others.the client uploads straight to storage —the API server never touches the bytes
Transcoding fans out per rendition and joins at the packager — which is why it is a DAG, not a queue.

The API and data model

Text
POST /v1/videos            → {video_id, upload_urls[], chunk_size}PUT  <pre-signed url>      → uploads one chunk directly to storagePOST /v1/videos/{id}/completeGET  /v1/videos/{id}       → metadata + manifest URL

The metadata store holds video identifier, uploader, title, description, status (uploading → processing → ready → failed), duration, and per-rendition manifest references. It is small — a few kilobytes per video, so a million uploads a day is a few gigabytes a day — and relational is a fine choice.