Course Content
Python Essentials for AI Engineer
6 sections · 48 lessons
How do you read a file line by line?
What you need to know
| Call | Returns | Memory |
|---|---|---|
f.read() | the whole file as one str | whole file |
f.readlines() | a list of every line | whole file, plus list overhead |
f.readline() | the next single line ("" at the end) | one line |
for line in f: | one line per loop | about one buffer |
Python
1with open("sample.log", "w", encoding="utf-8") as f:2 f.write("model=a status=200\nmodel=b status=429\nmodel=a status=500") # no final newline34with open("sample.log", encoding="utf-8") as f:5 for line_no, line in enumerate(f, start=1):6 print(line_no, repr(line))7# 1 'model=a status=200\n'8# 2 'model=b status=429\n'9# 3 'model=a status=500'Note the last line has no \n. That is why rstrip("\n") is safer than slicing off the last character. Use strip() only if leading and trailing spaces also do not matter.
Handy patterns
- Last N lines:
collections.deque(f, maxlen=100)keeps only the last 100 lines while streaming. - CSV:
csv.reader(f)andcsv.DictReader(f)stream rows and handle quoted commas correctly. - JSON Lines:
json.loads(line)for each line. - pandas:
pd.read_csv(path, chunksize=100_000)yields DataFrames of 100,000 rows at a time.
A real-life example
Your LLM gateway writes one log line per request, and the file is 8 GB. You need error counts per model. Loading it with readlines() would need more than 8 GB of RAM. Streaming keeps memory flat:
Python
1from collections import Counter23with open("gateway.log", "w", encoding="utf-8") as f: # a tiny stand-in for the real log4 f.write("model=chat-small status=200\nmodel=chat-large status=429\n"5 "model=chat-small status=500\nmodel=chat-large status=429\n")67errors = Counter()8with open("gateway.log", encoding="utf-8") as f:9 for line in f:10 fields = dict(part.split("=") for part in line.split())11 if int(fields["status"]) >= 400:12 errors[(fields["model"], fields["status"])] += 113print(errors.most_common())14# [(('chat-large', '429'), 2), (('chat-small', '500'), 1)]The same loop works on 8 KB or 8 GB. The answer — chat-large is being rate-limited — takes one pass and a few kilobytes of memory.
Follow-up questions to expect
- "What is the difference between
readline()andreadlines()?" —readline()returns the next single line;readlines()returns a list of all remaining lines. - "How would you read a 50 GB CSV?" — Stream it with
csv.reader, or usepandas.read_csv(..., chunksize=...), or a tool built for it like Polars or DuckDB. - "How do you get the last 10 lines of a big file?" —
deque(f, maxlen=10), which streams the file but keeps only 10 lines.