Course Content
System Design Interview
31 sections · 71 lessons
Hotel Reservation: requirements, scale and the inventory data model
"Design a hotel reservation system." The first problem in this course where a wrong answer is not slow or stale — it is two people arriving at the same room.
This lesson scopes the problem, sizes it, and makes the one modelling decision that carries the whole design. The second lesson builds the fast search path, the correct booking path, and the saga that ties booking to payment.
Why this bridges into Part V
Everything up to Section 23 could tolerate being briefly wrong. A feed missing a post for ten seconds is fine. A metric dropped is fine. A double-sold room is a person standing at a reception desk at midnight with a confirmation email and nowhere to sleep, and it costs real money to fix.
That does not make the whole system strongly consistent. It makes one small part of it strongly consistent, and the design's job is to keep that part as small as possible so everything else can stay fast. That framing — an island of strict correctness inside a mostly-relaxed system — is the bridge to Sections 28 to 30.
The questions that shape everything after
- How many hotels and rooms? This decides whether the inventory data is large or surprisingly small, and the arithmetic below shows it is the second.
- Do we sell individual rooms or room types? Guests book "a deluxe king", not room 412. Assigning the specific room at check-in rather than at booking removes an enormous amount of complexity, and it is what hotels actually do.
- Is pricing dynamic? If prices change by the minute, the price shown in search may not be the price at checkout, and you need a price-quote mechanism with an expiry.
- Is payment in scope? If yes, this becomes a two-service consistency problem, which Consistency across services and follow-ups covers.
- What is the cancellation policy? Free cancellation means inventory returns and overbooking becomes attractive.
- Is overbooking allowed? Ask this one explicitly. Real hotels deliberately sell more rooms than they have, because a predictable percentage of guests do not arrive. A candidate who assumes overbooking is a bug has not understood the business.
The assumptions this section uses
| Question | Assumption |
|---|---|
| Scale | 5,000 hotels, averaging 200 rooms, about 1 million rooms |
| Granularity | Room types, not individual rooms; specific room assigned at check-in |
| Pricing | Dynamic, with a quoted price held for 15 minutes |
| Payment | In scope, as a separate service |
| Cancellation | Free until 24 hours before arrival |
| Overbooking | Allowed, at a configurable percentage per hotel |
Requirements and scale
With those answers, the requirements are short. The numbers are where this problem surprises people.
Functional requirements
- Search hotels by location and date range, with availability and price.
- View a hotel's room types, photos, amenities, and rates.
- Reserve a room type for a date range, with payment.
- View, modify, and cancel a reservation.
- Hotel-side: manage inventory and rates.
Non-functional requirements
- No double-selling beyond the configured overbooking allowance. Hard.
- Search fast and highly available; stale by seconds is acceptable.
- Booking correct and durable; slower is acceptable.
- Reservations never lost, and a retried request never creates two.
The arithmetic
Inventory rows. 5,000 hotels × 20 room types × 730 days of forward booking window = 73 million rows. Each row is small — hotel ID, room type ID, date, total inventory, booked count, price — call it 60 bytes, so about 4.4 GB.
Say that number out loud, because it changes the conversation. The entire bookable inventory of a large chain fits comfortably in one database, and in memory. This is not a sharding problem. Candidates who reflexively shard here are solving a problem they do not have.
Booking rate. 1 million rooms at 70% occupancy with an average stay of 3 nights gives 1,000,000 × 0.7 ÷ 3 ≈ 233,000 bookings per day, which is 2.3 per second average. A seasonal peak of 10× is 23 bookings per second.
Search rate. People browse far more than they book. Assume 1,000 searches per booking — a ratio hotels and travel sites recognise. That is 2,300 searches per second average, and 23,000 per second at peak.
Reservation storage. 233,000 bookings/day × 365 = 85 million per year, at roughly 500 bytes each = 42 GB per year. Trivial.
What this rules in and out
- Sharding is not needed for inventory. One well-provisioned relational database with read replicas handles 23 writes per second and holds 4.4 GB. Say so, and say you would revisit if the business grew 100×.
- A relational database is the right choice for the booking path, because the correctness guarantee you need — atomic check-and-decrement across a date range — is exactly what transactions provide, and at 23 writes per second you will never approach its limits.
- The search path must not touch that database. 23,000 searches per second against the transactional store would generate lock contention against bookings for no benefit.
The data model
One modelling decision carries the whole design, and it is not obvious.
The wrong model: a row per room per booking
Store a rooms table and a reservations table, and answer "is a deluxe king free from the 3rd to the 6th?" by joining them and checking for overlaps. This works and gets slower every year, because every availability question becomes a range-overlap query over a growing reservations table, and the overlap predicate does not index well.
Worse, it makes the correctness question harder: to book, you must prove no overlapping reservation exists for any room of that type, which is a query whose result can be invalidated by a concurrent write. You are back to locking large ranges.
The right model: inventory per room type per date
Store one row for each combination of hotel, room type, and calendar date:
room_type_inventory hotel_id BIGINT room_type_id BIGINT date DATE total_inventory INT -- rooms of this type that exist total_reserved INT -- rooms of this type already sold for this date PRIMARY KEY (hotel_id, room_type_id, date)Availability becomes arithmetic on a single row: available = total_inventory − total_reserved. Overbooking becomes one multiplier: available = total_inventory × 1.1 − total_reserved, with the 1.1 configurable per hotel.
A three-night stay touches three rows. Booking is: in one transaction, increment total_reserved on all three, with a check constraint that it never exceeds the allowance. If any one of the three fails, the transaction rolls back and the guest is told the stay is not available — which is correct, because a partial stay is not what they asked for.
Supporting tables
hotelandroom_type— descriptive data, small, cached everywhere, changes rarely.rate— price per hotel, room type, and date, updated by the revenue system. Separated from inventory because it changes on a different schedule and is read by search far more often.reservation— one row per booking, carrying the guest, date range, room type, price, status, and an idempotency key. This is the record of truth for the guest.
Keeping rate out of the inventory row matters: search reads rates constantly, and mixing a hot-read column into the row that bookings write would create contention for no reason.