Python Essentials for AI Engineer

Course Content

Python Essentials for AI Engineer

6 sections · 48 lessons

What are Dunder (Magic) Methods?


What your syntax really callslen(ds)ds.__len__()ds[0]ds.__getitem__(0)a + ba.__add__(b)model(x)model.__call__(x)with f:f.__enter__()You writePython calls
A PyTorch-style Dataset is just the first two rows, which is why a plain class with them already works in a for loop.

What you need to know

Syntax to method

You writePython calls
Model(...)__new__, then __init__
repr(x), the REPLx.__repr__()
str(x), print(x)x.__str__() (falls back to __repr__)
len(x)x.__len__()
x[i]x.__getitem__(i)
for item in xx.__iter__()
a + b, a == ba.__add__(b), a.__eq__(b)
x(...)x.__call__(...)
with x:x.__enter__() and x.__exit__(...)

A small vector class

Python
class Vec:    def __init__(self, *values):        self.values = list(values)    def __repr__(self):        return f"Vec{tuple(self.values)}"    def __len__(self):        return len(self.values)    def __add__(self, other):        if not isinstance(other, Vec):            return NotImplemented          # let Python try other.__radd__ or raise TypeError        return Vec(*(a + b for a, b in zip(self.values, other.values)))    def __eq__(self, other):        return isinstance(other, Vec) and self.values == other.valuesa, b = Vec(1, 2), Vec(10, 20)print(a + b, len(a), a == Vec(1, 2))   # Vec(11, 22) 2 Trueprint(Vec.__hash__)                    # None -> defining __eq__ removed hashing

Return NotImplemented (not False or an exception) when an operation doesn't support the other type; Python then tries the other operand and finally raises a clear TypeError.

__repr__ vs __str__

__repr__ is for developers: unambiguous, ideally looks like the code to rebuild the object. __str__ is for end users. If you write only one, write __repr__.

A real-life example

PyTorch's Dataset contract is pure dunder methods: a dataset is any object with __len__ and __getitem__, and the DataLoader uses them to shuffle and batch. Here is the same idea for a fine-tuning set of support tickets, in plain Python:

Python
class TicketDataset:    def __init__(self, rows):        self.rows = rows    def __len__(self):        return len(self.rows)    def __getitem__(self, i):        text, label = self.rows[i]        return {"text": text.strip().lower(), "label": label}    def __repr__(self):        return f"TicketDataset(n={len(self)})"ds = TicketDataset([(" Refund pending ", 1), ("Change address", 0), ("UPI failed ", 1)])print(ds, len(ds))           # TicketDataset(n=3) 3print(ds[0])                 # {'text': 'refund pending', 'label': 1}print([row["label"] for row in ds])   # [1, 0, 1]

The last line works even though there is no __iter__: when a class has __getitem__, Python iterates by calling it with 0, 1, 2, ... until IndexError. Likewise, model(x) in PyTorch works because nn.Module defines __call__, which runs hooks and then your forward.

Follow-up questions to expect

  • "What is the difference between __str__ and __repr__?" — __repr__ is the developer view used in the REPL and logs; __str__ is the user-facing text used by print. str() falls back to __repr__.
  • "What does __call__ do?" — It makes instances callable like functions, which is how model(x) works in PyTorch.
  • "Why did my objects become unhashable after adding __eq__?" — Python sets __hash__ to None when you define __eq__, because equal objects must have equal hashes. Define __hash__ too, or use @dataclass(frozen=True).