Python Essentials for AI Engineer

Course Content

Python Essentials for AI Engineer

6 sections · 48 lessons

What is the __init__ method?


What you need to know

What happens on Dataset("train.csv")

  1. Python calls Dataset.__new__(Dataset, "train.csv"), which creates an empty instance.
  2. Python calls instance.__init__("train.csv"), which fills in the attributes.
  3. The call returns the instance.

You almost never write __new__ yourself. You write __init__.

Python
class Dataset:    def __init__(self, path, batch_size=32):        if batch_size <= 0:            raise ValueError(f"batch_size must be positive, got {batch_size}")        self.path = path        self.batch_size = batch_sizeds = Dataset("train.csv")print(ds.path, ds.batch_size)          # train.csv 32try:    Dataset("train.csv", batch_size=0)except ValueError as e:    print(e)                           # batch_size must be positive, got 0

Validating in __init__ means an invalid object can never exist, so the rest of the class can trust its own attributes.

It must return None

Python
class Broken:    def __init__(self):        return 42try:    Broken()except TypeError as e:    print(e)                           # __init__() should return None, not 'int'

Subclasses and super().__init__

If a child class defines its own __init__, the parent's __init__ does not run automatically. Call super().__init__(...), or the parent's attributes will be missing.

Alternative constructors with @classmethod

A classmethod receives the class (cls) instead of an instance and can build instances in different ways — from_csv_row, from_config, from_path. This keeps __init__ simple and puts slow work where a test can avoid it.

A real-life example

A RAG service's retriever originally loaded a 2 GB vector index inside __init__. Every unit test that created a Retriever spent 40 seconds loading it, so the team stopped writing tests. The refactor passes the index in, and moves loading to a classmethod:

Python
class Retriever:    def __init__(self, index, top_k=5):        self.index = index               # any object with a .search() method        self.top_k = top_k    @classmethod    def from_path(cls, path, top_k=5):        index = f"<index loaded from {path}>"   # real code: a slow load from disk        return cls(index, top_k)    def search(self, query):        return self.index.search(query, self.top_k)class FakeIndex:    def search(self, query, k):        return [f"doc-{i}" for i in range(k)]r = Retriever(FakeIndex(), top_k=2)      # tests: instant, no diskprint(r.search("refund policy"))         # ['doc-0', 'doc-1']prod = Retriever.from_path("/models/faq.index")   # production pathprint(prod.index)                        # <index loaded from /models/faq.index>

Tests now run in milliseconds with a fake index, and production code still gets a one-line constructor.

Follow-up questions to expect

  • "Is __init__ a constructor?" — Informally yes, but technically __new__ constructs the object and __init__ initialises it.
  • "What happens if __init__ returns a value?" — Python raises TypeError: __init__() should return None.
  • "Can a class have several __init__ methods?" — No; a second definition replaces the first. Use default arguments or @classmethod alternative constructors.