Python for AI and Data Science

Variables, Data Types, and Operators


You have a text file of temperature readings from a weather sensor. Three lines: 21.5, 19.8, 23.1. You read them in, add them up, and print the total. Python prints 21.519.823.1.

Nothing crashed. No error appeared. Python did exactly what you asked — it just did not do what you meant. You handed it three pieces of text that happened to look like numbers, and + on text means "stick these together end to end". The bug is silent, and it is the single most common way beginner data code produces confidently wrong answers.

Everything in this lesson exists to stop that happening. A variable is a label. The thing the label points at has a type, and the type decides what every operator does. Get the types right and the operators behave. Get them wrong and Python cheerfully computes nonsense.

Why the sensor totals came out as 21.519.823.1Read a linefrom the fileThe value isthe str "21.5"Plus joinsstrings, itdoes not addfloat()convertseach oneTotal is 64.4Nothing raised an error, because joining two strings is a perfectly legal operation.
The operator did exactly what the type asked for — the bug is the type, not the arithmetic.

A variable is a label, not a box

Most people picture a variable as a box you put a value into. That picture is wrong in Python, and the wrongness causes real bugs later. In Python, x = 5 does two separate things: it creates an object with the value 5 somewhere in memory, and it sticks the name x on that object like a luggage tag.

Python
temperature = 21.5      # the name 'temperature' now tags a float objectreading = temperature   # a SECOND tag on the SAME objecttemperature = 19.8      # move the 'temperature' tag to a NEW objectprint(reading)          # 21.5  -- 'reading' never moved

Because names are tags, you can retag freely. Python is dynamically typed: a name is not locked to one type, so x = 5 followed by x = "hello" is legal. That flexibility is why Python feels quick to write, and also why the silent-string bug above is possible.

Assignment never copies a value. It only adds another name to something that already exists.

Names must start with a letter or underscore, may contain letters, digits and underscores, and are case sensitive — Score and score are different names. The community convention is snake_case for variables and functions, CAPITALS for constants you do not intend to change.

The types you will actually meet

Whole numbers: int

Python integers have no size limit. They grow until your machine runs out of memory, which is unusual among programming languages and occasionally very handy.

Python
count = 42big = 2 ** 200        # 200-bit number, no overflow, no warningprint(len(str(big)))  # 61 digits

Decimals: float

Floats are stored in binary, and most decimal fractions have no exact binary form — the same way one third has no exact decimal form. This produces the result that surprises everyone once:

Python
print(0.1 + 0.2)              # 0.30000000000000004print(0.1 + 0.2 == 0.3)       # False

This is not a Python defect; it is how binary floating point works everywhere. The practical consequence: never compare floats with ==. Compare with a tolerance instead.

Python
import mathprint(math.isclose(0.1 + 0.2, 0.3))   # True

If you write if predicted_price == 100.0: in a data pipeline, it will occasionally miss a row that is off by 0.0000000000001, and you will spend an afternoon hunting a row that "should have matched".

Text: str

Strings are sequences of characters in single, double or triple quotes. They are immutable — you cannot change a character in place; every "modification" builds a new string.

Python
name = "ada lovelace"print(name.title())          # Ada Lovelaceprint(name.upper())          # ADA LOVELACEprint(name.split())          # ['ada', 'lovelace']print(name.replace("a", "@"))# @d@ lovel@ceprint(name)                  # ada lovelace  -- unchanged# f-strings: the modern way to build text from valuesscore = 0.8734print(f"{name.title()} scored {score:.1%}")   # Ada Lovelace scored 87.3%

That :.1% is a format specification: show as a percentage to one decimal place. Format specs save an enormous amount of rounding code when you print results.

True and false: bool

True and False are, underneath, the integers 1 and 0. This is not trivia — it is the trick behind counting filtered rows in nearly every data library.

Python
flags = [True, False, True, True]print(sum(flags))    # 3  -- counts the Trues

The four collections

Almost all data work lives in these four. The difference between them is the single most useful thing to memorise in this lesson.

TypeWritten asOrderedChangeableDuplicatesReach for it when
list[1, 2, 3]YesYesYesA sequence you will grow, sort or edit
tuple(1, 2, 3)YesNoYesA fixed record: coordinates, an array shape, a row
dict{"a": 1}Insertion orderYesKeys uniqueLookup by name; configuration; JSON
set{1, 2, 3}NoYesNoMembership tests and de-duplication
Python
scores = [88, 92, 79]scores.append(95)              # [88, 92, 79, 95]print(scores[0], scores[-1])   # 88 95   -- negative indexes count from the endprint(scores[1:3])             # [92, 79] -- slice: start included, stop excludedshape = (1000, 28, 28)         # a tuple: this will never change, so lock itrows, height, width = shape    # unpackingmodel = {"name": "rf", "depth": 10}print(model.get("seed", 42))   # 42 -- .get avoids a KeyError on a missing keymodel["seed"] = 7labels = {"cat", "dog", "cat"} # {'cat', 'dog'} -- duplicate silently dropped

The set is quietly the performance hero. Checking x in some_list scans every element; checking x in some_set jumps straight there. On a list of a million IDs, that is the difference between a loop that finishes in milliseconds and one that finishes over lunch.

Operators, and where they surprise you

Arithmetic

OperatorMeaningExampleResult
+ - *Add, subtract, multiply7 * 321
/True division — always a float10 / 25.0
//Floor division — rounds down-7 // 2-4
%Remainder17 % 52
**Power2 ** 101024

Two traps live in that table. First, / gives a float even when the division is exact, so 10 / 2 is 5.0, and using it as a list index raises TypeError. Use // for indexes. Second, // rounds towards negative infinity, not towards zero: -7 // 2 is -4, not -3. If you are splitting signed data into buckets, that off-by-one will shift a whole category.

% is more useful than it looks. i % 2 == 0 tests for even; i % 1000 == 0 prints progress every thousandth row without flooding your terminal.

Comparison and the chaining shortcut

The six comparisons — ==, !=, <, >, <=, >= — return booleans. Python lets you chain them the way mathematics does, which most languages do not:

Python
age = 25print(18 <= age < 65)     # True -- reads like maths, and is evaluated as maths

Logical operators short-circuit

and stops as soon as it hits something false; or stops as soon as it hits something true. That is not an optimisation detail you can ignore — it is a safety mechanism you should deliberately exploit:

Python
values = []# Safe: the length check fails, so the second half never runsif len(values) > 0 and values[0] > 10:    print("big first value")# Unsafe: swap the order and you get IndexError on an empty list

is versus == — the classic mistake

== asks "do these have the same value?". is asks "are these literally the same object in memory?". Beginners reach for is because it reads like English, and it works by accident on small numbers because Python caches them.

Python
a = [1, 2, 3]b = [1, 2, 3]print(a == b)    # True  -- same contentsprint(a is b)    # False -- two separate list objectsx = 256y = 256print(x is y)    # True  -- small ints are cached; this is an implementation detailx = 1000y = 1000print(x is y)    # False when typed line by line in the shell,                 # but True when run as a script -- the answer                 # depends on how Python compiled the code

That last result is the real warning. The same two lines give different answers depending on whether you type them into the shell or run them from a file, because the interpreter is free to reuse one object for equal constants. Code whose answer depends on that is broken even when it happens to print the right thing.

The rule: use is only for None, True and False. Everything else uses ==.

Use is None, never == None; and never is for numbers or strings, however well it seems to work in the shell.

Membership and compound assignment

Python
features = ["age", "income", "region"]print("age" in features)          # Trueprint("gender" not in features)   # Truetotal = 0total += 10       # same as total = total + 10total *= 3        # 30

Converting between types — and the two things that break

Back to the opening bug. The fix is one function call:

Python
raw = ["21.5", "19.8", "23.1"]print(sum(raw))                       # TypeError -- can't sum strings with 0readings = [float(v) for v in raw]    # [21.5, 19.8, 23.1]print(sum(readings))                  # 64.4
CallDoesFails when
int("42")Text to whole numberText is "42.0", "4,200", "" or "N/A"
int(9.99)Truncates to 9Never — but it truncates, it does not round
float("3.14")Text to decimalText has a currency symbol or stray space
str(42)Anything to textNever
list("abc")['a','b','c']Never — but it splits per character, which surprises people

The first failure mode is int() truncates rather than rounds: int(9.99) is 9. If you convert predicted ages that way, you systematically bias every value downwards. Use round() when you mean rounding.

The second is that real data is dirty. float("N/A"), float("") and float("1,250") all raise ValueError — and in a file of 100,000 rows, one bad cell stops the whole import. Convert defensively:

Python
def to_float(value, default=None):    try:        return float(value)    except (ValueError, TypeError):        return defaultprint(to_float("21.5"))   # 21.5print(to_float("N/A"))    # None

Mutability: the trap that costs the most debugging time

Lists, dictionaries and sets are mutable — they can be changed in place. Numbers, strings and tuples are immutable — they cannot. Combine mutability with the fact that assignment only copies a label, and you get this:

Python
original = [1, 2, 3]backup = original        # NOT a copy -- a second tag on the same listbackup.append(999)print(original)          # [1, 2, 3, 999]  -- your "backup" edited the original

This is where people get it wrong, and the reason is that the code looks like a copy. It is not. To actually copy:

Python
backup = original.copy()      # or list(original), or original[:]backup.append(999)print(original)               # [1, 2, 3]  -- safe# For nested structures, a shallow copy is not enough:import copygrid = [[0, 0], [0, 0]]shallow = grid.copy()shallow[0].append(1)print(grid)                   # [[0, 0, 1], [0, 0]]  -- inner lists still shareddeep = copy.deepcopy(grid)    # copies every level

The same trap appears with strings, but harmlessly, because strings are immutable — s.upper() cannot damage the original, so there is nothing to defend against.

If a function receives a list and modifies it, the caller's list changes too. Either document that clearly or copy on entry.

What this buys you when you build something

When you load a real dataset, the very first thing worth doing is checking what you actually got — not what the file appeared to contain. A three-line habit prevents most silent-wrongness bugs:

Python
import csvwith open("readings.csv") as f:    rows = list(csv.DictReader(f))sample = rows[0]for key, value in sample.items():    print(f"{key:<12} {value!r:>12}   {type(value).__name__}")

Every value from a CSV arrives as str. Every value from JSON arrives already typed. Every value from a database driver depends on the driver. Until you have looked, you do not know — and + on the wrong type will not tell you.

Three habits carry most of the weight. Convert types explicitly at the boundary where data enters your program, so the rest of your code can assume clean types. Compare floats with math.isclose, never ==. And whenever you assign one mutable object to a second name, ask yourself whether you meant a copy — because Python assumed you did not.

Check your understanding

0 of 3 answered

1.You read the values "4" and "5" from a CSV file and compute a + b. What do you get?

2.backup = original is followed by backup.append(99), where original is a list. What happens to original?

3.Which of these is the safe way to check whether a model's predicted value equals 0.3?