Course Content
Python for AI and Data Science
5 sections · 13 lessons
Variables, Data Types, and Operators
You have a text file of temperature readings from a weather sensor. Three lines: 21.5, 19.8, 23.1. You read them in, add them up, and print the total. Python prints 21.519.823.1.
Nothing crashed. No error appeared. Python did exactly what you asked — it just did not do what you meant. You handed it three pieces of text that happened to look like numbers, and + on text means "stick these together end to end". The bug is silent, and it is the single most common way beginner data code produces confidently wrong answers.
Everything in this lesson exists to stop that happening. A variable is a label. The thing the label points at has a type, and the type decides what every operator does. Get the types right and the operators behave. Get them wrong and Python cheerfully computes nonsense.
A variable is a label, not a box
Most people picture a variable as a box you put a value into. That picture is wrong in Python, and the wrongness causes real bugs later. In Python, x = 5 does two separate things: it creates an object with the value 5 somewhere in memory, and it sticks the name x on that object like a luggage tag.
1temperature = 21.5 # the name 'temperature' now tags a float object2reading = temperature # a SECOND tag on the SAME object3temperature = 19.8 # move the 'temperature' tag to a NEW object45print(reading) # 21.5 -- 'reading' never movedBecause names are tags, you can retag freely. Python is dynamically typed: a name is not locked to one type, so x = 5 followed by x = "hello" is legal. That flexibility is why Python feels quick to write, and also why the silent-string bug above is possible.
Assignment never copies a value. It only adds another name to something that already exists.
Names must start with a letter or underscore, may contain letters, digits and underscores, and are case sensitive — Score and score are different names. The community convention is snake_case for variables and functions, CAPITALS for constants you do not intend to change.
The types you will actually meet
Whole numbers: int
Python integers have no size limit. They grow until your machine runs out of memory, which is unusual among programming languages and occasionally very handy.
count = 42big = 2 ** 200 # 200-bit number, no overflow, no warningprint(len(str(big))) # 61 digitsDecimals: float
Floats are stored in binary, and most decimal fractions have no exact binary form — the same way one third has no exact decimal form. This produces the result that surprises everyone once:
print(0.1 + 0.2) # 0.30000000000000004print(0.1 + 0.2 == 0.3) # FalseThis is not a Python defect; it is how binary floating point works everywhere. The practical consequence: never compare floats with ==. Compare with a tolerance instead.
import mathprint(math.isclose(0.1 + 0.2, 0.3)) # TrueIf you write if predicted_price == 100.0: in a data pipeline, it will occasionally miss a row that is off by 0.0000000000001, and you will spend an afternoon hunting a row that "should have matched".
Text: str
Strings are sequences of characters in single, double or triple quotes. They are immutable — you cannot change a character in place; every "modification" builds a new string.
1name = "ada lovelace"2print(name.title()) # Ada Lovelace3print(name.upper()) # ADA LOVELACE4print(name.split()) # ['ada', 'lovelace']5print(name.replace("a", "@"))# @d@ lovel@ce6print(name) # ada lovelace -- unchanged78# f-strings: the modern way to build text from values9score = 0.873410print(f"{name.title()} scored {score:.1%}") # Ada Lovelace scored 87.3%That :.1% is a format specification: show as a percentage to one decimal place. Format specs save an enormous amount of rounding code when you print results.
True and false: bool
True and False are, underneath, the integers 1 and 0. This is not trivia — it is the trick behind counting filtered rows in nearly every data library.
flags = [True, False, True, True]print(sum(flags)) # 3 -- counts the TruesThe four collections
Almost all data work lives in these four. The difference between them is the single most useful thing to memorise in this lesson.
| Type | Written as | Ordered | Changeable | Duplicates | Reach for it when |
|---|---|---|---|---|---|
list | [1, 2, 3] | Yes | Yes | Yes | A sequence you will grow, sort or edit |
tuple | (1, 2, 3) | Yes | No | Yes | A fixed record: coordinates, an array shape, a row |
dict | {"a": 1} | Insertion order | Yes | Keys unique | Lookup by name; configuration; JSON |
set | {1, 2, 3} | No | Yes | No | Membership tests and de-duplication |
1scores = [88, 92, 79]2scores.append(95) # [88, 92, 79, 95]3print(scores[0], scores[-1]) # 88 95 -- negative indexes count from the end4print(scores[1:3]) # [92, 79] -- slice: start included, stop excluded56shape = (1000, 28, 28) # a tuple: this will never change, so lock it7rows, height, width = shape # unpacking89model = {"name": "rf", "depth": 10}10print(model.get("seed", 42)) # 42 -- .get avoids a KeyError on a missing key11model["seed"] = 71213labels = {"cat", "dog", "cat"} # {'cat', 'dog'} -- duplicate silently droppedThe set is quietly the performance hero. Checking x in some_list scans every element; checking x in some_set jumps straight there. On a list of a million IDs, that is the difference between a loop that finishes in milliseconds and one that finishes over lunch.
Operators, and where they surprise you
Arithmetic
| Operator | Meaning | Example | Result |
|---|---|---|---|
+ - * | Add, subtract, multiply | 7 * 3 | 21 |
/ | True division — always a float | 10 / 2 | 5.0 |
// | Floor division — rounds down | -7 // 2 | -4 |
% | Remainder | 17 % 5 | 2 |
** | Power | 2 ** 10 | 1024 |
Two traps live in that table. First, / gives a float even when the division is exact, so 10 / 2 is 5.0, and using it as a list index raises TypeError. Use // for indexes. Second, // rounds towards negative infinity, not towards zero: -7 // 2 is -4, not -3. If you are splitting signed data into buckets, that off-by-one will shift a whole category.
% is more useful than it looks. i % 2 == 0 tests for even; i % 1000 == 0 prints progress every thousandth row without flooding your terminal.
Comparison and the chaining shortcut
The six comparisons — ==, !=, <, >, <=, >= — return booleans. Python lets you chain them the way mathematics does, which most languages do not:
age = 25print(18 <= age < 65) # True -- reads like maths, and is evaluated as mathsLogical operators short-circuit
and stops as soon as it hits something false; or stops as soon as it hits something true. That is not an optimisation detail you can ignore — it is a safety mechanism you should deliberately exploit:
1values = []2# Safe: the length check fails, so the second half never runs3if len(values) > 0 and values[0] > 10:4 print("big first value")56# Unsafe: swap the order and you get IndexError on an empty listis versus == — the classic mistake
== asks "do these have the same value?". is asks "are these literally the same object in memory?". Beginners reach for is because it reads like English, and it works by accident on small numbers because Python caches them.
1a = [1, 2, 3]2b = [1, 2, 3]3print(a == b) # True -- same contents4print(a is b) # False -- two separate list objects56x = 2567y = 2568print(x is y) # True -- small ints are cached; this is an implementation detail910x = 100011y = 100012print(x is y) # False when typed line by line in the shell,13 # but True when run as a script -- the answer14 # depends on how Python compiled the codeThat last result is the real warning. The same two lines give different answers depending on whether you type them into the shell or run them from a file, because the interpreter is free to reuse one object for equal constants. Code whose answer depends on that is broken even when it happens to print the right thing.
The rule: use is only for None, True and False. Everything else uses ==.
Use
is None, never== None; and neverisfor numbers or strings, however well it seems to work in the shell.
Membership and compound assignment
1features = ["age", "income", "region"]2print("age" in features) # True3print("gender" not in features) # True45total = 06total += 10 # same as total = total + 107total *= 3 # 30Converting between types — and the two things that break
Back to the opening bug. The fix is one function call:
1raw = ["21.5", "19.8", "23.1"]23print(sum(raw)) # TypeError -- can't sum strings with 04readings = [float(v) for v in raw] # [21.5, 19.8, 23.1]5print(sum(readings)) # 64.4| Call | Does | Fails when |
|---|---|---|
int("42") | Text to whole number | Text is "42.0", "4,200", "" or "N/A" |
int(9.99) | Truncates to 9 | Never — but it truncates, it does not round |
float("3.14") | Text to decimal | Text has a currency symbol or stray space |
str(42) | Anything to text | Never |
list("abc") | ['a','b','c'] | Never — but it splits per character, which surprises people |
The first failure mode is int() truncates rather than rounds: int(9.99) is 9. If you convert predicted ages that way, you systematically bias every value downwards. Use round() when you mean rounding.
The second is that real data is dirty. float("N/A"), float("") and float("1,250") all raise ValueError — and in a file of 100,000 rows, one bad cell stops the whole import. Convert defensively:
1def to_float(value, default=None):2 try:3 return float(value)4 except (ValueError, TypeError):5 return default67print(to_float("21.5")) # 21.58print(to_float("N/A")) # NoneMutability: the trap that costs the most debugging time
Lists, dictionaries and sets are mutable — they can be changed in place. Numbers, strings and tuples are immutable — they cannot. Combine mutability with the fact that assignment only copies a label, and you get this:
1original = [1, 2, 3]2backup = original # NOT a copy -- a second tag on the same list3backup.append(999)4print(original) # [1, 2, 3, 999] -- your "backup" edited the originalThis is where people get it wrong, and the reason is that the code looks like a copy. It is not. To actually copy:
1backup = original.copy() # or list(original), or original[:]2backup.append(999)3print(original) # [1, 2, 3] -- safe45# For nested structures, a shallow copy is not enough:6import copy7grid = [[0, 0], [0, 0]]8shallow = grid.copy()9shallow[0].append(1)10print(grid) # [[0, 0, 1], [0, 0]] -- inner lists still shared1112deep = copy.deepcopy(grid) # copies every levelThe same trap appears with strings, but harmlessly, because strings are immutable — s.upper() cannot damage the original, so there is nothing to defend against.
If a function receives a list and modifies it, the caller's list changes too. Either document that clearly or copy on entry.
What this buys you when you build something
When you load a real dataset, the very first thing worth doing is checking what you actually got — not what the file appeared to contain. A three-line habit prevents most silent-wrongness bugs:
1import csv23with open("readings.csv") as f:4 rows = list(csv.DictReader(f))56sample = rows[0]7for key, value in sample.items():8 print(f"{key:<12} {value!r:>12} {type(value).__name__}")Every value from a CSV arrives as str. Every value from JSON arrives already typed. Every value from a database driver depends on the driver. Until you have looked, you do not know — and + on the wrong type will not tell you.
Three habits carry most of the weight. Convert types explicitly at the boundary where data enters your program, so the rest of your code can assume clean types. Compare floats with math.isclose, never ==. And whenever you assign one mutable object to a second name, ask yourself whether you meant a copy — because Python assumed you did not.
Check your understanding
0 of 3 answered
1.You read the values "4" and "5" from a CSV file and compute a + b. What do you get?
2.backup = original is followed by backup.append(99), where original is a list. What happens to original?
3.Which of these is the safe way to check whether a model's predicted value equals 0.3?