Python Essentials for AI Engineer

Course Content

Python Essentials for AI Engineer

6 sections · 48 lessons

What is a Dictionary?


One lookup: stats["Swiggy"]hash("Swiggy")gives an integerThe integerpicks a slotin the tableStored keychecked with ==Value returned, noother keys scannedA collision only means probing another candidate slot.
The work does not grow with the number of keys, so 10 merchants or 10 million cost the same per lookup.

What you need to know

The basic operations

Python
user = {"name": "Asha", "role": "ML Engineer"}print(user["role"])              # ML Engineerprint(user.get("city", "NA"))    # NA  -> no KeyError for a missing keyuser["city"] = "Pune"            # insert or updatedel user["role"]                 # deleteprint("city" in user)            # True  -> checks KEYS, not valuesprint(user)                      # {'name': 'Asha', 'city': 'Pune'}print(user | {"team": "search"}) # {'name': 'Asha', 'city': 'Pune', 'team': 'search'}

The | merge operator creates a new dict (Python 3.9+). Writing to an existing key replaces the value, which is why keys are unique.

How a hash table gives O(1)

When you write user["city"], Python computes hash("city"), uses it to pick a slot in an internal array, checks that the key stored there really equals "city", and returns the value. It does not scan the other keys. With a list of pairs you would compare key after key — O(n). With a dict the time stays about the same for 10 keys or 10 million.

Two keys can land in the same slot (a collision); Python then probes other slots. Collisions are rare with a good hash, so the average stays O(1), though the theoretical worst case is O(n).

Why keys must be hashable

If a key could change after insertion, its hash would change and Python would look in the wrong slot. So keys must be immutable types: str, int, float, tuple of hashables, frozenset. A list or dict cannot be a key.

Ordering

Since Python 3.7 the language guarantees that iteration follows insertion order. Older answers saying "dicts are unordered" are out of date.

A real-life example

You have a day's transactions and want the count and total per merchant — the core of any spending dashboard:

Python
txns = [("Swiggy", 349), ("Zomato", 512), ("Swiggy", 180), ("IRCTC", 1450), ("Swiggy", 99)]stats = {}for merchant, amount in txns:    entry = stats.get(merchant, {"count": 0, "total": 0})    entry["count"] += 1    entry["total"] += amount    stats[merchant] = entryfor merchant, s in stats.items():    print(merchant, s)# Swiggy {'count': 3, 'total': 628}# Zomato {'count': 1, 'total': 512}# IRCTC {'count': 1, 'total': 1450}

Each row costs one O(1) lookup, so 5 rows or 5 million rows scale linearly. The same shape powers LLM calls: messages = [{"role": "system", "content": "..."}, {"role": "user", "content": "..."}] is a list of dicts.

Follow-up questions to expect

  • "What happens on a hash collision?" — Python probes for another free slot and compares keys with == to find the right one. Lookups stay O(1) on average.
  • "dict vs defaultdict vs Counter?" — defaultdict(list) creates a default value for missing keys automatically; Counter is a dict specialised for counting, with helpers like most_common(3).
  • "Can two keys be 1 and 1.0?" — No. 1 == 1.0 and they hash the same, so they are the same key.