Python Essentials for AI Engineer

Course Content

Python Essentials for AI Engineer

6 sections · 48 lessons

What is the difference between iterable and iterator?


Streaming a 5 GB JSONL file in batches of 64File objectyields one lineread_jsonlyields one textbatched collects64, then yieldsembed(batch),then thebatch is freedMemory holds one batch whether the file has 150 lines or 150 million.
Each stage pulls the next item only when asked, so nothing downstream forces the whole file into memory.

What you need to know

The protocol by hand

Python
docs = ["doc1", "doc2"]          # an ITERABLEit = iter(docs)                  # ask it for an ITERATORprint(next(it))                  # doc1print(next(it))                  # doc2try:    next(it)except StopIteration:    print("exhausted")           # exhaustedprint(iter(it) is it)            # True -> an iterator is its own iterator

A for loop does exactly this: it calls iter(docs), calls next() until StopIteration, and hides the exception.

Exhaustion

Python
gen = (x * x for x in range(3))  # a generator is an iteratorprint(list(gen))                 # [0, 1, 4]print(list(gen))                 # []  -> already used up, no error

Generators: iterators you write with yield

A function containing yield returns a generator. Each next() runs the function until the next yield, hands out that value, and pauses with all local variables kept. Nothing is computed until someone asks, so a generator can describe a million rows while holding only one.

A real-life example

You must embed a 5 GB JSONL file (one JSON document per line) with an API that accepts at most 64 texts per request. Loading the file into a list would need more RAM than the container has. Two small generators stream it instead:

Python
import io, jsondef read_jsonl(f):    for line in f:                        # a file object is an iterator of lines        if line.strip():            yield json.loads(line)["text"]def batched(items, size):    batch = []    for item in items:        batch.append(item)        if len(batch) == size:            yield batch            batch = []    if batch:        yield batch                       # the last, smaller batchf = io.StringIO("".join(json.dumps({"text": f"doc {i}"}) + "\n" for i in range(150)))for batch in batched(read_jsonl(f), 64):    print(len(batch), batch[0])           # embed(batch) would go here# 64 doc 0# 64 doc 64# 22 doc 128

At any moment only one batch of 64 texts is in memory, whether the file has 150 lines or 150 million. Python 3.12 added itertools.batched, which does the same job (it yields tuples); the hand-written version above works on 3.11 too.

Follow-up questions to expect

  • "Is a list an iterator?" — No. It is an iterable: next([1, 2]) raises TypeError. iter([1, 2]) gives an iterator.
  • "Why can I loop over a list twice but not a generator?" — Each loop over a list creates a new iterator; a generator is the iterator, and once exhausted it stays empty.
  • "What is the difference between yield and return?" — return ends the function; yield hands out one value and pauses, so the function can continue on the next next() call.