Course Content
Python Essentials for AI Engineer
6 sections · 48 lessons
How do you work with JSON files?
What you need to know
Load and dump
1import json23cfg = {"model": "chat-small", "temperature": 0.2, "stop": None, "tags": ("faq", "hi")}4with open("cfg.json", "w", encoding="utf-8") as f:5 json.dump(cfg, f, indent=2, ensure_ascii=False)6with open("cfg.json", encoding="utf-8") as f:7 back = json.load(f)8print(back) # {'model': 'chat-small', 'temperature': 0.2, 'stop': None, 'tags': ['faq', 'hi']}9print(json.dumps({1: "a"})) # {"1": "a"} -> keys become strings10print(json.dumps({"msg": "नमस्ते"}, ensure_ascii=False)) # {"msg": "नमस्ते"}indent=2makes files readable for humans; leave it out for compact payloads.ensure_ascii=Falsekeeps Hindi, Tamil or emoji readable instead ofन...escapes.- The round trip is not perfect: the tuple came back as a list, and the int key
1would come back as the string"1".
Types JSON cannot hold
1import json2from datetime import datetime3try:4 json.dumps({"at": datetime(2026, 9, 24, 10, 30)})5except TypeError as e:6 print(e) # Object of type datetime is not JSON serializable7print(json.dumps({"at": datetime(2026, 9, 24, 10, 30)}, default=str))8# {"at": "2026-09-24 10:30:00"}default= is a function called for any object JSON does not understand. Convert sets with list(), Decimal with str(), and NumPy arrays with .tolist().
JSON Lines
A normal JSON file holds one big value, so you must load all of it. JSON Lines (.jsonl) puts one JSON object on each line. You can stream it line by line, append new records without rewriting, and a single corrupt line does not ruin the file. Most fine-tuning and batch APIs use it.
A real-life example
You ask an LLM to return {"label": ..., "confidence": ...}. Most replies parse, but some arrive wrapped in Markdown code fences or with extra text, and json.loads raises json.JSONDecodeError. A small, defensive parser handles the common cases and reports the rest:
1import json23def parse_reply(text):4 text = text.strip()5 if text.startswith("```"):6 text = text.strip("`").removeprefix("json").strip() # drop ```json ... ``` fences7 try:8 data = json.loads(text)9 except json.JSONDecodeError as e:10 return None, f"not JSON: {e.msg}"11 if not {"label", "confidence"} <= data.keys():12 return None, "missing keys"13 return data, None1415print(parse_reply('{"label": "refund", "confidence": 0.91}'))16# ({'label': 'refund', 'confidence': 0.91}, None)17print(parse_reply('```json\n{"label": "kyc", "confidence": 0.7}\n```'))18# ({'label': 'kyc', 'confidence': 0.7}, None)19print(parse_reply("Sure! The label is refund."))20# (None, 'not JSON: Expecting value')In production, a None result triggers one retry with the error message included in the prompt. Better still, use your provider's structured-output or JSON mode, or validate with Pydantic — but you still need this fallback for the replies that slip through.
Follow-up questions to expect
- "What is the difference between
loadandloads?" —loadreads from a file object;loadsparses a string. The same goes fordumpanddumps. - "How do you serialise a
datetime?" — Convert it to an ISO string with.isoformat(), or passdefault=strtojson.dumps. - "Why use JSON Lines for datasets?" — Each line is independent, so you can stream, append and recover from a bad line without loading everything.