Python for AI and Data Science

File Handling and Modules


A script runs overnight, processes 200,000 records, prints "done", and writes them to results.csv. In the morning the file is 8 kilobytes and stops mid-row. Nothing crashed. The output is simply incomplete.

The cause: the script opened the file, wrote to it, and never closed it. Writes do not go straight to disk — they collect in a memory buffer and get flushed when the buffer fills or the file is closed. The process ended, the buffer was discarded, and the last chunk of data evaporated.

Python
f = open("results.csv", "w")     # the seed of the problemfor row in rows:    f.write(row)# no f.close() -- and if anything above raises, you never reach it anyway

The whole of file handling in Python is arranged around making sure this cannot happen, and the mechanism is a single keyword.

Where 200,000 records became 8 kilobytesopen()hands backa file objectWritesaccumulatein a bufferBuffer flushesonly when it fillsScript printsdone and exitsThe last partialbuffer is lostThe with statement closes the file on the way out, flush included, even if the block raises.
Nothing crashed because nothing failed — the rows were simply still in memory when the process ended.

The with statement is not optional

Python
with open("results.csv", "w") as f:    for row in rows:        f.write(row)# file is closed and flushed here -- even if the loop raised an exception

with creates a block that guarantees cleanup. When the block ends — normally, by return, or because an exception blew through it — Python closes the file. There is no code path where the buffer is lost.

Every file you open should be opened with with. There is no situation in ordinary data work where the bare open is the better choice.

Three ways to read, and when each is right

Python
with open("data.txt") as f:    whole = f.read()          # one big string -- entire file in memorywith open("data.txt") as f:    lines = f.readlines()     # list of strings, each ending in "\n"with open("data.txt") as f:    for line in f:            # one line at a time -- constant memory        process(line.strip())
MethodMemory usedUse forBreaks on
f.read()Whole fileSmall config or text you must search acrossA 5 GB log — your process dies
f.readlines()Whole file, plus list overheadWhen you need random access to line 40,000Same, slightly worse
for line in fOne lineAnything large; the default choiceNothing — but you get one pass only

Iterating the file object directly is lazy: Python reads a block at a time and hands you lines as you ask for them. That is why it works on files bigger than your RAM.

Modes: the second argument to open

ModeMeaningIf the file existsIf it does not
"r"Read (the default)Opens itFileNotFoundError
"w"WriteErases it immediatelyCreates it
"a"AppendAdds at the endCreates it
"x"Exclusive createFileExistsErrorCreates it
"rb" / "wb"BinaryAs above, no text decodingAs above

Mode "w" truncates the file the instant it opens, before you write a single byte. If you open your only copy of a dataset in "w" by mistake, it is gone — not overwritten at the end, gone at the start. When you are appending to a results log, "a" is what you want; when you want a hard error rather than a silent overwrite, "x" is the safe choice.

Encoding: the error that only appears on someone else's machine

A text file is bytes. Turning bytes into characters requires knowing the encoding, and Python's default depends on the operating system. Your script reads a file happily on Linux and then a colleague on Windows gets:

Text
UnicodeDecodeError: 'charmap' codec can't decode byte 0x9d in position 3244

Always state the encoding. UTF-8 is the correct answer roughly always:

Python
with open("data.csv", encoding="utf-8") as f:    ...# For files with the odd corrupt byte you cannot fix at source:with open("messy.csv", encoding="utf-8", errors="replace") as f:    ...    # bad bytes become a placeholder character instead of raising

Files exported from Excel often carry a byte-order mark at the start, which turns your first column name into "id" and makes row["id"] raise KeyError. Use encoding="utf-8-sig" and it disappears.

CSV: rows and columns as text

You can split lines on commas yourself. You should not. A single quoted field containing a comma — "Smith, John" — breaks a naive split, and it is guaranteed to appear eventually in real data.

Python
import csvwith open("people.csv", newline="", encoding="utf-8") as f:    reader = csv.DictReader(f)      # first row becomes the keys    for row in reader:        print(row["name"], row["age"])   # both are STRINGSwith open("out.csv", "w", newline="", encoding="utf-8") as f:    writer = csv.DictWriter(f, fieldnames=["name", "age"])    writer.writeheader()    writer.writerows([{"name": "Ada", "age": 36}])

Two details in that code are easy to skip and both cause real bugs. newline="" stops Python from translating line endings, which on Windows otherwise produces a blank row between every data row. And every value from a CSV arrives as a string — row["age"] is "36", not 36, so summing an age column concatenates text unless you convert.

JSON: nested structures, preserved

CSV is flat. The moment a record contains a list or another record, CSV stops being able to express it, and JSON takes over. JSON maps almost one-to-one onto Python's own types.

Python
import jsonconfig = {    "model": "random_forest",    "params": {"depth": 10, "trees": 200},    "features": ["age", "income"],    "active": True,}with open("config.json", "w", encoding="utf-8") as f:    json.dump(config, f, indent=2)      # indent makes it human-readablewith open("config.json", encoding="utf-8") as f:    loaded = json.load(f)print(loaded["params"]["depth"])        # 10 -- still an int, not a string
JSONPython
objectdict
arraylist
stringstr
numberint or float
true / falseTrue / False
nullNone

Note what is missing from that table: tuples, sets, dates and NumPy numbers. json.dump raises TypeError: Object of type ndarray is not JSON serializable the first time you try to save model output directly. Convert to plain lists and floats first.

json.dump writes to a file; json.dumps returns a string. The trailing s stands for "string", and mixing the two up is a rite of passage.

Which format to use

CSVJSON
ShapeFlat table onlyArbitrary nesting
TypesEverything is textNumbers, booleans and null preserved
SizeCompactLarger — keys repeat on every record
StreamingRow by row, triviallyWhole document at once, normally
Best forDatasets, spreadsheets, model inputConfig, API payloads, nested records

Handling failure instead of crashing

File operations fail for reasons outside your control: the file moved, the disk filled, permissions changed, the content is malformed. A pipeline that dies on the first bad file is a pipeline someone has to babysit.

Python
import jsondef load_config(path, fallback=None):    try:        with open(path, encoding="utf-8") as f:            return json.load(f)    except FileNotFoundError:        print(f"{path} not found; using defaults")        return fallback or {}    except json.JSONDecodeError as e:        print(f"{path} is not valid JSON (line {e.lineno}); using defaults")        return fallback or {}    except PermissionError:        raise                        # this one you DO want to stop for

Catch the specific exceptions you know how to recover from. A bare except: swallows everything, including your own typos and the interrupt when you press Ctrl-C, and turns a five-second fix into an afternoon.

Paths that work everywhere

Hard-coding "data\\raw\\file.csv" breaks on Linux; "data/raw/file.csv" is fine on both but string-joining paths gets ugly fast. pathlib handles it:

Python
from pathlib import Pathbase = Path("data") / "raw"target = base / "readings.csv"print(target.exists(), target.suffix, target.stem)base.mkdir(parents=True, exist_ok=True)     # create the folder tree if neededfor csv_file in base.glob("*.csv"):         # every CSV in the folder    print(csv_file.name)text = target.read_text(encoding="utf-8")   # shortcut for small files

Modules: code that lives somewhere else

A module is a .py file. A package is a folder of them. import runs that file once and binds its names so you can use them.

Python
import math                       # math.sqrt(16)import numpy as np                # np.array(...)  -- alias, universal conventionfrom pathlib import Path          # Path directly in your namespacefrom collections import Counter, defaultdict

The one form to avoid is from module import *. It dumps every public name into your file, so two libraries that both define load silently overwrite each other and you get the wrong function with no indication which. Whichever import ran last wins, which means the bug moves when you reorder your imports.

The standard library is bigger than people expect

ModuleGives youTypical one-liner
mathNumeric functionsmath.isclose(a, b)
randomSampling and shufflingrandom.sample(rows, 100)
statisticsMean, median, stdevstatistics.median(vals)
datetimeDates and durationsdatetime.strptime(s, "%Y-%m-%d")
collectionsCounter, defaultdictCounter(labels).most_common(3)
itertoolsLazy combinatoricsitertools.combinations(cols, 2)
os / pathlibFilesystemos.environ["API_KEY"]
loggingStructured outputlogging.warning("%d dropped", n)

Counter and defaultdict deserve special mention because they delete so much boilerplate:

Python
from collections import Counter, defaultdictlabels = ["cat", "dog", "cat", "bird", "cat"]print(Counter(labels).most_common(2))    # [('cat', 3), ('dog', 1)]groups = defaultdict(list)               # missing keys spring into existencefor word in ["apple", "avocado", "banana"]:    groups[word[0]].append(word)print(dict(groups))                      # {'a': ['apple', 'avocado'], 'b': ['banana']}

Writing your own module

Put reusable functions in their own file and import them. That is the entire mechanism.

Python
# dataprep.pydef load_rows(path):    import csv    with open(path, newline="", encoding="utf-8-sig") as f:        return list(csv.DictReader(f))def to_float(value, default=None):    try:        return float(value)    except (TypeError, ValueError):        return defaultif __name__ == "__main__":    print("quick self-test:", to_float("3.5"), to_float("n/a"))
Python
# analysis.pyfrom dataprep import load_rows, to_floatrows = load_rows("data/raw/readings.csv")values = [to_float(r["temp"]) for r in rows]

What if __name__ == "__main__": actually does

Python sets a variable called __name__ in every module. When you run a file directly it is "__main__"; when the file is imported it is the module's name. So that block runs on python dataprep.py and is skipped on import dataprep.

Leave it out and every import of your module executes your test code, your file writes and your plots. This is exactly what goes wrong when someone imports a script that had a long training run at the bottom of the file: importing it starts training.

A module should define things when imported and do things only when run directly. The __main__ guard is what enforces that boundary.

Packages

A folder becomes a package when you organise modules inside it:

Text
project/├── data/│   ├── raw/│   └── processed/├── src/│   ├── __init__.py│   ├── dataprep.py│   └── models.py├── notebooks/└── requirements.txt
Python
from src.dataprep import load_rows

The __init__.py file marks the folder as a package and runs when the package is first imported; it can be empty. Keeping raw data separate from processed data is not fussiness — it means a broken cleaning step can always be re-run from an original you have not overwritten.

What this looks like in a working project

Put together, the pattern for getting data off disk safely is short and worth making a habit:

Python
import csv, json, loggingfrom pathlib import Pathlogging.basicConfig(level=logging.INFO, format="%(levelname)s %(message)s")def load_dataset(path, numeric_fields=()):    """Read a CSV into dicts, converting the named fields to floats."""    path = Path(path)    if not path.exists():        raise FileNotFoundError(f"no dataset at {path.resolve()}")    rows, dropped = [], 0    with path.open(newline="", encoding="utf-8-sig") as f:        for row in csv.DictReader(f):            try:                for field in numeric_fields:                    row[field] = float(row[field])            except (ValueError, KeyError):                dropped += 1                continue            rows.append(row)    logging.info("%s: kept %d rows, dropped %d", path.name, len(rows), dropped)    return rowsdef save_report(summary, path):    path = Path(path)    path.parent.mkdir(parents=True, exist_ok=True)    with path.open("w", encoding="utf-8") as f:        json.dump(summary, f, indent=2)if __name__ == "__main__":    rows = load_dataset("data/raw/readings.csv", numeric_fields=["temp", "humidity"])    save_report({"n_rows": len(rows)}, "data/processed/summary.json")

Every failure mode from this lesson is defended against there. with guarantees the flush, so no truncated output. The explicit encoding means it behaves the same on every machine. Type conversion happens at the boundary, so nothing downstream has to wonder whether temp is a number. Bad rows are counted and skipped rather than crashing the run — and crucially, they are logged, so if 40% of your file silently vanished you find out immediately rather than after training a model on a quarter of the data.