Python Essentials for AI Engineer

Course Content

Python Essentials for AI Engineer

6 sections · 48 lessons

How do you read a file line by line?


What you need to know

CallReturnsMemory
f.read()the whole file as one strwhole file
f.readlines()a list of every linewhole file, plus list overhead
f.readline()the next single line ("" at the end)one line
for line in f:one line per loopabout one buffer
Python
with open("sample.log", "w", encoding="utf-8") as f:    f.write("model=a status=200\nmodel=b status=429\nmodel=a status=500")   # no final newlinewith open("sample.log", encoding="utf-8") as f:    for line_no, line in enumerate(f, start=1):        print(line_no, repr(line))# 1 'model=a status=200\n'# 2 'model=b status=429\n'# 3 'model=a status=500'

Note the last line has no \n. That is why rstrip("\n") is safer than slicing off the last character. Use strip() only if leading and trailing spaces also do not matter.

Handy patterns

  • Last N lines: collections.deque(f, maxlen=100) keeps only the last 100 lines while streaming.
  • CSV: csv.reader(f) and csv.DictReader(f) stream rows and handle quoted commas correctly.
  • JSON Lines: json.loads(line) for each line.
  • pandas: pd.read_csv(path, chunksize=100_000) yields DataFrames of 100,000 rows at a time.

A real-life example

Your LLM gateway writes one log line per request, and the file is 8 GB. You need error counts per model. Loading it with readlines() would need more than 8 GB of RAM. Streaming keeps memory flat:

Python
from collections import Counterwith open("gateway.log", "w", encoding="utf-8") as f:     # a tiny stand-in for the real log    f.write("model=chat-small status=200\nmodel=chat-large status=429\n"            "model=chat-small status=500\nmodel=chat-large status=429\n")errors = Counter()with open("gateway.log", encoding="utf-8") as f:    for line in f:        fields = dict(part.split("=") for part in line.split())        if int(fields["status"]) >= 400:            errors[(fields["model"], fields["status"])] += 1print(errors.most_common())# [(('chat-large', '429'), 2), (('chat-small', '500'), 1)]

The same loop works on 8 KB or 8 GB. The answer — chat-large is being rate-limited — takes one pass and a few kilobytes of memory.

Follow-up questions to expect

  • "What is the difference between readline() and readlines()?" — readline() returns the next single line; readlines() returns a list of all remaining lines.
  • "How would you read a 50 GB CSV?" — Stream it with csv.reader, or use pandas.read_csv(..., chunksize=...), or a tool built for it like Polars or DuckDB.
  • "How do you get the last 10 lines of a big file?" — deque(f, maxlen=10), which streams the file but keeps only 10 lines.