Course Content
Python for AI and Data Science
5 sections · 13 lessons
File Handling and Modules
A script runs overnight, processes 200,000 records, prints "done", and writes them to results.csv. In the morning the file is 8 kilobytes and stops mid-row. Nothing crashed. The output is simply incomplete.
The cause: the script opened the file, wrote to it, and never closed it. Writes do not go straight to disk — they collect in a memory buffer and get flushed when the buffer fills or the file is closed. The process ended, the buffer was discarded, and the last chunk of data evaporated.
1f = open("results.csv", "w") # the seed of the problem2for row in rows:3 f.write(row)4# no f.close() -- and if anything above raises, you never reach it anywayThe whole of file handling in Python is arranged around making sure this cannot happen, and the mechanism is a single keyword.
The with statement is not optional
1with open("results.csv", "w") as f:2 for row in rows:3 f.write(row)4# file is closed and flushed here -- even if the loop raised an exceptionwith creates a block that guarantees cleanup. When the block ends — normally, by return, or because an exception blew through it — Python closes the file. There is no code path where the buffer is lost.
Every file you open should be opened with
with. There is no situation in ordinary data work where the bareopenis the better choice.
Three ways to read, and when each is right
1with open("data.txt") as f:2 whole = f.read() # one big string -- entire file in memory34with open("data.txt") as f:5 lines = f.readlines() # list of strings, each ending in "\n"67with open("data.txt") as f:8 for line in f: # one line at a time -- constant memory9 process(line.strip())| Method | Memory used | Use for | Breaks on |
|---|---|---|---|
f.read() | Whole file | Small config or text you must search across | A 5 GB log — your process dies |
f.readlines() | Whole file, plus list overhead | When you need random access to line 40,000 | Same, slightly worse |
for line in f | One line | Anything large; the default choice | Nothing — but you get one pass only |
Iterating the file object directly is lazy: Python reads a block at a time and hands you lines as you ask for them. That is why it works on files bigger than your RAM.
Modes: the second argument to open
| Mode | Meaning | If the file exists | If it does not |
|---|---|---|---|
"r" | Read (the default) | Opens it | FileNotFoundError |
"w" | Write | Erases it immediately | Creates it |
"a" | Append | Adds at the end | Creates it |
"x" | Exclusive create | FileExistsError | Creates it |
"rb" / "wb" | Binary | As above, no text decoding | As above |
Mode "w" truncates the file the instant it opens, before you write a single byte. If you open your only copy of a dataset in "w" by mistake, it is gone — not overwritten at the end, gone at the start. When you are appending to a results log, "a" is what you want; when you want a hard error rather than a silent overwrite, "x" is the safe choice.
Encoding: the error that only appears on someone else's machine
A text file is bytes. Turning bytes into characters requires knowing the encoding, and Python's default depends on the operating system. Your script reads a file happily on Linux and then a colleague on Windows gets:
UnicodeDecodeError: 'charmap' codec can't decode byte 0x9d in position 3244Always state the encoding. UTF-8 is the correct answer roughly always:
1with open("data.csv", encoding="utf-8") as f:2 ...34# For files with the odd corrupt byte you cannot fix at source:5with open("messy.csv", encoding="utf-8", errors="replace") as f:6 ... # bad bytes become a placeholder character instead of raisingFiles exported from Excel often carry a byte-order mark at the start, which turns your first column name into "id" and makes row["id"] raise KeyError. Use encoding="utf-8-sig" and it disappears.
CSV: rows and columns as text
You can split lines on commas yourself. You should not. A single quoted field containing a comma — "Smith, John" — breaks a naive split, and it is guaranteed to appear eventually in real data.
1import csv23with open("people.csv", newline="", encoding="utf-8") as f:4 reader = csv.DictReader(f) # first row becomes the keys5 for row in reader:6 print(row["name"], row["age"]) # both are STRINGS78with open("out.csv", "w", newline="", encoding="utf-8") as f:9 writer = csv.DictWriter(f, fieldnames=["name", "age"])10 writer.writeheader()11 writer.writerows([{"name": "Ada", "age": 36}])Two details in that code are easy to skip and both cause real bugs. newline="" stops Python from translating line endings, which on Windows otherwise produces a blank row between every data row. And every value from a CSV arrives as a string — row["age"] is "36", not 36, so summing an age column concatenates text unless you convert.
JSON: nested structures, preserved
CSV is flat. The moment a record contains a list or another record, CSV stops being able to express it, and JSON takes over. JSON maps almost one-to-one onto Python's own types.
1import json23config = {4 "model": "random_forest",5 "params": {"depth": 10, "trees": 200},6 "features": ["age", "income"],7 "active": True,8}910with open("config.json", "w", encoding="utf-8") as f:11 json.dump(config, f, indent=2) # indent makes it human-readable1213with open("config.json", encoding="utf-8") as f:14 loaded = json.load(f)1516print(loaded["params"]["depth"]) # 10 -- still an int, not a string| JSON | Python |
|---|---|
object | dict |
array | list |
string | str |
number | int or float |
true / false | True / False |
null | None |
Note what is missing from that table: tuples, sets, dates and NumPy numbers. json.dump raises TypeError: Object of type ndarray is not JSON serializable the first time you try to save model output directly. Convert to plain lists and floats first.
json.dump writes to a file; json.dumps returns a string. The trailing s stands for "string", and mixing the two up is a rite of passage.
Which format to use
| CSV | JSON | |
|---|---|---|
| Shape | Flat table only | Arbitrary nesting |
| Types | Everything is text | Numbers, booleans and null preserved |
| Size | Compact | Larger — keys repeat on every record |
| Streaming | Row by row, trivially | Whole document at once, normally |
| Best for | Datasets, spreadsheets, model input | Config, API payloads, nested records |
Handling failure instead of crashing
File operations fail for reasons outside your control: the file moved, the disk filled, permissions changed, the content is malformed. A pipeline that dies on the first bad file is a pipeline someone has to babysit.
1import json23def load_config(path, fallback=None):4 try:5 with open(path, encoding="utf-8") as f:6 return json.load(f)7 except FileNotFoundError:8 print(f"{path} not found; using defaults")9 return fallback or {}10 except json.JSONDecodeError as e:11 print(f"{path} is not valid JSON (line {e.lineno}); using defaults")12 return fallback or {}13 except PermissionError:14 raise # this one you DO want to stop forCatch the specific exceptions you know how to recover from. A bare except: swallows everything, including your own typos and the interrupt when you press Ctrl-C, and turns a five-second fix into an afternoon.
Paths that work everywhere
Hard-coding "data\\raw\\file.csv" breaks on Linux; "data/raw/file.csv" is fine on both but string-joining paths gets ugly fast. pathlib handles it:
1from pathlib import Path23base = Path("data") / "raw"4target = base / "readings.csv"56print(target.exists(), target.suffix, target.stem)7base.mkdir(parents=True, exist_ok=True) # create the folder tree if needed89for csv_file in base.glob("*.csv"): # every CSV in the folder10 print(csv_file.name)1112text = target.read_text(encoding="utf-8") # shortcut for small filesModules: code that lives somewhere else
A module is a .py file. A package is a folder of them. import runs that file once and binds its names so you can use them.
1import math # math.sqrt(16)2import numpy as np # np.array(...) -- alias, universal convention3from pathlib import Path # Path directly in your namespace4from collections import Counter, defaultdictThe one form to avoid is from module import *. It dumps every public name into your file, so two libraries that both define load silently overwrite each other and you get the wrong function with no indication which. Whichever import ran last wins, which means the bug moves when you reorder your imports.
The standard library is bigger than people expect
| Module | Gives you | Typical one-liner |
|---|---|---|
math | Numeric functions | math.isclose(a, b) |
random | Sampling and shuffling | random.sample(rows, 100) |
statistics | Mean, median, stdev | statistics.median(vals) |
datetime | Dates and durations | datetime.strptime(s, "%Y-%m-%d") |
collections | Counter, defaultdict | Counter(labels).most_common(3) |
itertools | Lazy combinatorics | itertools.combinations(cols, 2) |
os / pathlib | Filesystem | os.environ["API_KEY"] |
logging | Structured output | logging.warning("%d dropped", n) |
Counter and defaultdict deserve special mention because they delete so much boilerplate:
1from collections import Counter, defaultdict23labels = ["cat", "dog", "cat", "bird", "cat"]4print(Counter(labels).most_common(2)) # [('cat', 3), ('dog', 1)]56groups = defaultdict(list) # missing keys spring into existence7for word in ["apple", "avocado", "banana"]:8 groups[word[0]].append(word)9print(dict(groups)) # {'a': ['apple', 'avocado'], 'b': ['banana']}Writing your own module
Put reusable functions in their own file and import them. That is the entire mechanism.
1# dataprep.py2def load_rows(path):3 import csv4 with open(path, newline="", encoding="utf-8-sig") as f:5 return list(csv.DictReader(f))67def to_float(value, default=None):8 try:9 return float(value)10 except (TypeError, ValueError):11 return default1213if __name__ == "__main__":14 print("quick self-test:", to_float("3.5"), to_float("n/a"))1# analysis.py2from dataprep import load_rows, to_float34rows = load_rows("data/raw/readings.csv")5values = [to_float(r["temp"]) for r in rows]What if __name__ == "__main__": actually does
Python sets a variable called __name__ in every module. When you run a file directly it is "__main__"; when the file is imported it is the module's name. So that block runs on python dataprep.py and is skipped on import dataprep.
Leave it out and every import of your module executes your test code, your file writes and your plots. This is exactly what goes wrong when someone imports a script that had a long training run at the bottom of the file: importing it starts training.
A module should define things when imported and do things only when run directly. The
__main__guard is what enforces that boundary.
Packages
A folder becomes a package when you organise modules inside it:
project/├── data/│ ├── raw/│ └── processed/├── src/│ ├── __init__.py│ ├── dataprep.py│ └── models.py├── notebooks/└── requirements.txtfrom src.dataprep import load_rowsThe __init__.py file marks the folder as a package and runs when the package is first imported; it can be empty. Keeping raw data separate from processed data is not fussiness — it means a broken cleaning step can always be re-run from an original you have not overwritten.
What this looks like in a working project
Put together, the pattern for getting data off disk safely is short and worth making a habit:
1import csv, json, logging2from pathlib import Path34logging.basicConfig(level=logging.INFO, format="%(levelname)s %(message)s")56def load_dataset(path, numeric_fields=()):7 """Read a CSV into dicts, converting the named fields to floats."""8 path = Path(path)9 if not path.exists():10 raise FileNotFoundError(f"no dataset at {path.resolve()}")1112 rows, dropped = [], 013 with path.open(newline="", encoding="utf-8-sig") as f:14 for row in csv.DictReader(f):15 try:16 for field in numeric_fields:17 row[field] = float(row[field])18 except (ValueError, KeyError):19 dropped += 120 continue21 rows.append(row)2223 logging.info("%s: kept %d rows, dropped %d", path.name, len(rows), dropped)24 return rows2526def save_report(summary, path):27 path = Path(path)28 path.parent.mkdir(parents=True, exist_ok=True)29 with path.open("w", encoding="utf-8") as f:30 json.dump(summary, f, indent=2)3132if __name__ == "__main__":33 rows = load_dataset("data/raw/readings.csv", numeric_fields=["temp", "humidity"])34 save_report({"n_rows": len(rows)}, "data/processed/summary.json")Every failure mode from this lesson is defended against there. with guarantees the flush, so no truncated output. The explicit encoding means it behaves the same on every machine. Type conversion happens at the boundary, so nothing downstream has to wonder whether temp is a number. Bad rows are counted and skipped rather than crashing the run — and crucially, they are logged, so if 40% of your file silently vanished you find out immediately rather than after training a model on a quarter of the data.