Python for AI and Data Science

Control Flow, Functions, and Comprehensions


Here is a task that sounds trivial. You have a list of sensor readings and you want to throw away the negative ones, which are known errors. Write the obvious loop:

Python
readings = [12, -1, -3, 8, 15]for r in readings:    if r < 0:        readings.remove(r)print(readings)   # [12, -3, 8, 15]

A negative number survived. No error, no warning — the list just came back wrong. What happened is that remove shifted every later element down by one position while the loop's internal counter kept marching forward, so the loop stepped straight over -3.

The fix is not a cleverer loop. It is to stop mutating and start building: clean = [r for r in readings if r >= 0]. One line, correct, and faster. This lesson is about the three tools that decide whether your data code is correct and readable — branching, functions, and comprehensions — and about the specific places each of them bites.

Removing items while looping skips the next one5-1-280123removed hereslidesleft, skippedThe index moves to 2 while -2 has already shifted down into slot 1.
A comprehension builds a new list instead of mutating the one being walked, which is why it cannot skip.

Branching: choosing a path

An if statement runs a block only when a condition is true. The colon and the indentation are the syntax — Python has no braces, so indentation is not style, it is the structure of your program.

Python
accuracy = 0.87if accuracy >= 0.90:    verdict = "ship it"elif accuracy >= 0.80:    verdict = "promising, needs tuning"else:    verdict = "back to the data"print(verdict)    # promising, needs tuning

Order matters and the first match wins. If you had written the >= 0.80 branch first, an accuracy of 0.95 would report "promising, needs tuning" — every value above 0.90 also passes the 0.80 test, so the stricter condition must come first. This is the most common branching bug: overlapping conditions in the wrong order, producing an answer that is plausible enough that nobody notices.

Truthiness, and what it costs you

Python lets you put almost anything where a condition is expected. Empty things are false; non-empty things are true.

ValueTreated as
0, 0.0False
"", [], {}, set()False
NoneFalse
-1, 0.001, "0", [0]True

So if rows: is the idiomatic way to say "if there is any data". But look at the last row of that table. The string "0" is true, and the list [0] is true, because they are not empty. And here is the trap that costs real money:

Python
def apply_discount(price, discount=None):    if not discount:                 # WRONG        discount = 0.10    return price * (1 - discount)print(apply_discount(100, 0))        # 90.0 -- a 0% discount became 10%

0 is falsy, so "no discount requested" was indistinguishable from "discount not supplied". The correct test asks the question you actually mean:

Python
if discount is None:    discount = 0.10

Use truthiness for emptiness. Use is None for absence. Confusing the two silently turns a legitimate zero into a missing value.

For short assignments, the conditional expression is compact and readable: label = "pass" if score >= 50 else "fail". Do not nest these; two levels deep and nobody can read it.

Loops: repeating work

Python has two. Choosing the wrong one produces code that works but reads badly, or code that never terminates.

forwhile
RunsOnce per item in a sequenceUntil a condition turns false
Use whenYou know what you are iterating overYou do not know how many rounds it takes
Typical jobEvery row, every file, every epochConverging on a value, retrying a request, reading until end of stream
Failure modeMutating the thing you are iteratingForgetting to change the condition — infinite loop
Python
for name in ["ada", "alan", "grace"]:    print(name.title())# enumerate gives you the index too -- do not maintain a counter by handfor i, name in enumerate(["ada", "alan", "grace"], start=1):    print(f"{i}. {name}")# zip walks two sequences in stepfor name, score in zip(["ada", "alan"], [98, 91]):    print(f"{name}: {score}")

A while loop that converges is the standard shape in numerical work — gradient descent is exactly this:

Python
value, step, iterations = 10.0, 0.1, 0while abs(value) > 0.001 and iterations < 1000:    value -= step * value        # move towards zero    iterations += 1print(f"{value:.4f} after {iterations} steps")

Notice the second condition. Without a hard iteration cap, a while loop that fails to converge — because of a bad step size, or a value that oscillates — runs forever. Always give a convergence loop an escape hatch.

break, continue, and the loop else

break exits the loop entirely. continue skips to the next iteration. There is also a feature almost unique to Python: a loop can have an else clause that runs only if the loop finished without hitting break.

Python
targets = [4, 9, 16, 25]for n in targets:    if n % 2 == 0 and n > 10:        print(f"found {n}")        breakelse:    print("no match in the whole list")   # runs only if break never fired

That saves the usual dance of setting a found = False flag and checking it afterwards.

Functions: the unit of reuse

Suppose you normalise three columns by writing the same four lines three times. When you later discover the formula should subtract the median rather than the mean, you must find and fix all three — and you will miss one. A function makes the change happen in one place.

Python
def normalise(values):    """Scale values to the range 0-1. Returns a new list."""    low, high = min(values), max(values)    if high == low:        return [0.0] * len(values)     # guard against dividing by zero    return [(v - low) / (high - low) for v in values]print(normalise([10, 20, 30]))         # [0.0, 0.5, 1.0]print(normalise([7, 7, 7]))            # [0.0, 0.0, 0.0]

That guard clause matters. A constant column — every value identical — appears in real data all the time, and without the check the function raises ZeroDivisionError halfway through a long pipeline.

Arguments: four kinds

Python
def train(data, epochs=10, *extra_sets, verbose=False, **options):    ...
FormNameBehaviour
dataPositionalRequired; matched by position
epochs=10DefaultOptional; used when the caller omits it
*extra_setsVariadic positionalCollects any surplus positional arguments into a tuple
**optionsVariadic keywordCollects any surplus named arguments into a dict

Calling with keywords rather than position is worth the extra typing. split(data, 0.2, True, 42) is unreadable at the call site; split(data, test_size=0.2, shuffle=True, seed=42) explains itself and survives a change in parameter order.

Returning several things at once

Python
def summarise(values):    n = len(values)    mean = sum(values) / n    spread = max(values) - min(values)    return n, mean, spread          # this is a tuplecount, avg, rng = summarise([4, 8, 15, 16, 23, 42])print(count, round(avg, 2), rng)    # 6 18.0 38

A function with no return statement returns None. That bites when you write a function that mutates a list and forget to return it — result = my_sort(data) then quietly assigns None, and the failure surfaces three lines later as TypeError: 'NoneType' object is not iterable.

The mutable default argument

This one deserves its own section because it is genuinely counter-intuitive and it appears in production code everywhere.

Python
def collect(item, bucket=[]):     # DANGEROUS    bucket.append(item)    return bucketprint(collect("a"))    # ['a']print(collect("b"))    # ['a', 'b']  -- where did 'a' come from?print(collect("c"))    # ['a', 'b', 'c']

Default values are evaluated once, when the def line runs — not on each call. So every call that omits bucket shares the same list object, and it accumulates forever. In a long-running service this is a slow memory leak with wrong answers attached.

Python
def collect(item, bucket=None):   # correct    if bucket is None:        bucket = []    bucket.append(item)    return bucket

Never use a list, dict or set as a default argument. Use None and build the real default inside the function body.

Scope: which name wins

Python resolves a name by looking in four places, in order: the Local function, any Enclosing function, the Global module level, and finally Python's Built-ins.

Python
threshold = 0.5              # globaldef classify(p):    threshold = 0.8          # local -- shadows the global, does NOT change it    return p > thresholdprint(classify(0.6))         # Falseprint(threshold)             # 0.5 -- untouched

Assigning inside a function always creates a local name unless you explicitly declare global. Reading works outward, writing works inward. The practical rule is to avoid globals altogether: pass what you need in as an argument, return what you produce. A function that depends on hidden module state is a function you cannot test in isolation.

A separate trap: never name a variable after a built-in. list = [1, 2, 3] works, and then list(some_tuple) fails with TypeError: 'list' object is not callable for the rest of the session. The same goes for sum, type, id, str, max and dict.

Comprehensions: build, do not mutate

A comprehension turns "make an empty list, loop, append" into one expression. It is faster than the loop — the interpreter skips a method lookup and call per item — and it forces you into the safe pattern of building a new collection rather than editing one in flight.

Python
values = [4, -1, 9, -6, 16]# the long waysquares = []for v in values:    if v > 0:        squares.append(v ** 2)# the same thingsquares = [v ** 2 for v in values if v > 0]     # [16, 81, 256]

Read it in the order it executes: for v in values, then if v > 0, then v ** 2. The output expression is written first but evaluated last.

A condition placed before the for is a different construct — a conditional expression — and it maps rather than filters:

Python
[v if v > 0 else 0 for v in values]    # [4, 0, 9, 0, 16]  -- same length, clipped[v for v in values if v > 0]           # [4, 9, 16]        -- shorter, filtered

Dicts and sets use the same syntax with different brackets:

Python
names = ["ada", "alan", "grace"]lengths = {n: len(n) for n in names}            # {'ada': 3, 'alan': 4, 'grace': 5}initials = {n[0] for n in names}                # {'a', 'g'}  -- set, so deduped# inverting a dict is a one-linerinverse = {v: k for k, v in lengths.items()}

Generator expressions, for when the data is large

Swap the square brackets for round ones and nothing is built in memory — values are produced one at a time as they are consumed.

Python
total = sum(v ** 2 for v in range(10_000_000))   # constant memorysquares = [v ** 2 for v in range(10_000_000)]    # ~400 MB list, then summed

If you only need to feed the values into sum, max, any or a loop, use a generator. If you need to index into the result or use it twice, you need the list — a generator is exhausted after one pass and silently yields nothing the second time.

When a comprehension is the wrong choice

Comprehensions win up to about the point where you need a second nested loop plus two conditions. Past that, they compress the logic into something nobody — including you, in a month — can debug.

Python
# unreadableresult = [transform(x, y) for x in rows if x.valid for y in x.cols if y > threshold and y != 0]

If you cannot read it aloud in one breath, write the loop. Clarity is worth more than compactness, and the speed difference is negligible next to the cost of a bug you cannot see.

Putting the three together

Here is the shape almost every real data-cleaning routine takes: a function with a guard clause, a comprehension doing the filtering, and branching to classify what survived.

Python
def clean_and_grade(raw, minimum=0.0, maximum=100.0):    """Drop unparseable and out-of-range scores, then grade what remains."""    if not raw:        return [], {"kept": 0, "dropped": 0}    numeric = []    for value in raw:        try:            numeric.append(float(value))        except (TypeError, ValueError):            continue                       # unparseable -- skip it    kept = [v for v in numeric if minimum <= v <= maximum]    graded = [        (v, "A" if v >= 90 else "B" if v >= 75 else "C" if v >= 60 else "F")        for v in kept    ]    return graded, {"kept": len(kept), "dropped": len(raw) - len(kept)}data = ["92", "88.5", "n/a", "150", "45", None, "61"]graded, report = clean_and_grade(data)print(graded)   # [(92.0, 'A'), (88.5, 'B'), (45.0, 'F'), (61.0, 'C')]print(report)   # {'kept': 4, 'dropped': 3}

Three things in that code are doing real defensive work. The empty-input guard means the function never divides by zero or indexes into nothing. The try inside the loop means one bad cell out of a million does not kill the run. And the comprehension builds a new list, so the caller's original data is untouched — which means if the result looks wrong, you still have the input to inspect.

That last point is the through-line. Branching decides, functions contain, comprehensions build. The bug at the top of this page happened because the code edited its input while walking it. Almost every version of that bug disappears the moment you make a new collection instead.