Course Content
Introduction to AI
4 sections · 10 lessons
AI Tools — ChatGPT, Copilot, Midjourney, and More
There are now thousands of AI tools, and the list changes every month. Memorising them is pointless — half of what you learned would be obsolete by the time you finished.
What does not go out of date is the underlying shape. Nearly every tool you will meet belongs to one of five families, and within a family they behave in fundamentally similar ways: the same strengths, the same failure modes, the same tricks for getting good results.
So this lesson teaches the families rather than the products. Learn these, and a tool released next year will be immediately legible to you.
Family 1: Conversational assistants
Chat interfaces on top of a large language model — ChatGPT, Claude, Gemini, and the many products built on their APIs. Current assistants are multimodal: besides typed text, they accept photos, screenshots, PDFs and speech, and many can answer out loud.
What is actually happening when you type
Your message is not "understood" in the way the interface implies. Here is the real sequence:
- Your text is broken into tokens — roughly word fragments. "Understanding" might become "under", "stand", "ing".
- The whole conversation so far is fed in as context.
- The model produces a probability for every possible next token.
- One is chosen, appended, and the process repeats from step 2.
The reply is generated one fragment at a time, each one conditioned on everything before it. There is no plan drawn up in advance and then written out. This single fact explains most of the behaviour that confuses people.
Three consequences worth internalising
It has no memory between conversations. Start a new chat and everything is gone. When a product appears to remember you, something outside the model is storing notes and quietly pasting them into the context each time. The model itself learns nothing from talking to you.
There is a hard limit on how much it can consider at once. This is the context window. Everything — your instructions, the conversation history, any document you paste — competes for the same space. Current models hold far more than early ones did, often hundreds of pages of text, but the space is still finite. When a conversation outgrows it, something has to give: older material is dropped or squeezed into a summary. This is why a long conversation can seem to "forget" what you agreed near the start.
It will state falsehoods with complete confidence. The model is optimising for plausible text, not true text. A fabricated citation has exactly the right shape — a believable author, a real-sounding journal, a plausible year. Nothing in the mechanism checks whether it exists.
The practical rule that follows: use these tools for things you can verify, or for things where being slightly wrong does not matter. Drafting, rephrasing, brainstorming, explaining a concept you will then check — excellent. Looking up a specific fact, statistic or legal citation you cannot verify — dangerous.
Getting better results
Most advice about "prompting" is folklore. These four things genuinely work, and each follows directly from how the model operates.
Supply the context it cannot have. The model does not know your situation. "Write a project update" produces generic filler. "Write a project update for a non-technical client, explaining a two-week delay caused by a third-party dependency, apologetic but not grovelling, under 150 words" produces something usable. You are not being polite by adding detail — you are narrowing the space of plausible continuations.
Show an example of what good looks like. One sample of the output format you want does more than three paragraphs describing it. The model is a pattern matcher; give it a pattern.
Ask for the reasoning before the answer. Because text is generated sequentially, tokens produced earlier influence what comes later. Asking it to work through a problem step by step before concluding measurably improves accuracy on anything involving logic or arithmetic — the intermediate steps become context for the final answer. Reasoning models, covered in the first section, do this working on their own before they reply, so with them you rarely need to ask.
Iterate rather than perfecting the first message. Get a rough result, say what is wrong with it, repeat. This is faster than trying to specify everything up front, and it plays to the format's strength.
Family 2: Coding assistants
GitHub Copilot, Cursor, Claude Code, and the coding modes of general assistants. Same underlying technology, trained heavily on source code.
They come in two shapes. The older one suggests: it completes the line or function you are typing. The newer one is a coding agent: you describe a change, and it reads the relevant files, edits several of them, runs the tests and tries again when they fail. The agent shape saves more time and raises the stakes, because it changes many lines you did not type.
Why code suits this technology unusually well
Code has properties that make it a better fit than most text:
- It is highly patterned. Vast amounts of real code are near-repetitions of well-known shapes.
- It is verifiable. You can run it. Unlike a paragraph of prose, wrongness is often immediately demonstrable.
- Enormous public training data exists, already organised into files, functions and repositories.
Where it helps most, and least
| Strong | Weak |
|---|---|
| Boilerplate and repetitive structure | Design decisions across a whole system |
| Writing tests for existing code | Team conventions it has not been shown or told about |
| Explaining unfamiliar code | Subtle concurrency and performance work |
| Translating between languages | Security-critical logic |
| Remembering syntax you rarely use | Novel algorithms with no precedent to draw on |
There is one failure mode specific to coding assistants and it deserves a warning, because it turns a wrong answer into a security hole.
The tool can invent a library function, or a whole package, that does not exist. It looks exactly right — plausible name, sensible arguments, matching the conventions of that library. It simply is not real. Security researchers have shown that an attacker can register a package under a name models commonly invent and wait for someone to install it.
The discipline that follows: read what it produces before accepting it, and with an agent, review the whole change before it is merged. Suggested code that you do not understand is a liability, not a shortcut. If you cannot explain what a line does, you cannot maintain it and you certainly cannot debug it at 2am.
Family 3: Image generation
Midjourney, Stable Diffusion, Google's Imagen, and the image generators built into general assistants such as ChatGPT and Gemini.
How diffusion works, briefly
Most image generators are built on a method called diffusion (some newer ones, built into chat assistants, work differently, but the behaviour below applies broadly). The mechanism is genuinely elegant and takes one paragraph to convey.
During training, the system takes real images and progressively adds random noise until nothing remains but static. It learns to undo each step — to look at a slightly noisy image and predict what the cleaner version looked like.
To generate something new, you start with pure noise and run that learned process in reverse, steered at every step by your text description. The image is not retrieved or collaged; it emerges from noise, guided toward whatever matches your words.
What this predicts about its behaviour
Understanding the mechanism explains the quirks:
- Text inside images can come out garbled. Early models learned that letter-like marks appear in certain contexts without ever learning spelling. Current leading models, trained with far more attention to text, usually get a short sign or headline right, but long passages and small print still slip.
- Hands and fine detail slip. Hands appear in enormous variety — different counts of visible fingers, angles, occlusions. Current models get them right far more often than early ones did, but an extra finger or a fused object still turns up, so check the details before you publish.
- The same prompt gives different images. You start from different random noise each time. This is inherent, not a fault.
- Style words are disproportionately powerful. "Watercolour", "cinematic lighting", "35mm film" pull hard, because they described large, visually consistent clusters in the training data.
Writing prompts that work
An effective image prompt usually names four things: subject, setting, style, and framing or light.
Weak: a dogBetter: a golden retriever puppy sitting in tall summer grass, late afternoon sunlight from behind, shallow depth of field, photographed on 35mm filmOne counter-intuitive point: negative instructions often fail. "A street with no cars" tends to produce cars, because the phrase makes cars strongly present in the description. Describe what you do want — "an empty pedestrianised street at dawn" — rather than what you do not.
Family 4: Audio and voice
Three distinct capabilities usually lumped together.
Speech to text. The most mature and reliable of the three. Modern systems handle accents and background noise well and are genuinely production-ready for transcription, captioning and voice commands.
Text to speech. Now close enough to human that most listeners do not notice in short passages. Used for audiobooks, accessibility and interfaces.
Voice cloning. A short sample is enough to reproduce someone's voice. This is where the serious ethical weight sits — it has enabled a wave of fraud in which a relative's cloned voice is used to request money urgently. If you build with this, consent and disclosure are not optional extras.
Family 5: Retrieval-augmented tools
This family answers the two biggest weaknesses of the others: no knowledge of your private documents, and confident invention.
The idea is to stop asking the model to know things, and instead give it the relevant material at the moment of the question.
1. You ask: "What is our refund policy for damaged goods?"2. The system searches your documents for relevant passages.3. It pastes those passages into the model's context, with your question.4. The model answers using the supplied text.5. The answer cites which document each part came from.Two things improve at once. The model can answer about material it was never trained on, and it is far less likely to invent, because the correct answer is sitting right there in its context. The citations also let a reader check.
This is the shape behind most serious business deployments — internal knowledge assistants, documentation search, customer support tools. If you build one useful AI product in your career, there is a good chance it will be this family.
When the families combine: agents
The five families are blurring. A current assistant can search the web, read a PDF you upload, write and run code to check a calculation, and produce an image, all in one conversation. The most capable setups go further and act as agents: given a goal, the model picks a tool, looks at the result, and repeats until the task is done — working through a website, fixing a failing test, or compiling a summary from twenty sources.
Everything in this lesson still applies to each step, and one thing gets more serious. A wrong chat reply can be ignored; a wrong agent may already have sent the email or deleted the file. Give an agent the narrowest permissions that do the job, and keep a human approval step before anything that is hard to undo.
Choosing a tool
| Need | Family | Watch out for |
|---|---|---|
| Draft, rewrite, explain | Conversational | Confident invention on facts |
| Write or understand code | Coding assistant | Imaginary functions; read before accepting |
| Create visuals | Image generation | Text and hands; licensing of output |
| Transcribe or narrate | Audio | Consent for any cloned voice |
| Answer from your own documents | Retrieval-augmented | Quality depends on the search step |
Things worth being careful about
What you paste may be retained. Consumer tiers of many services may use submitted content to improve their systems. Client data, personal data and unreleased work do not belong in a consumer chat box. Business tiers usually contract this away — read the terms rather than assuming.
Ownership of generated output is unsettled. Copyright status varies by jurisdiction and is actively being litigated. For anything commercial, check the specific terms of the specific tool rather than relying on general impressions.
Cost scales with use, not with seats. These services charge by volume of text processed. A feature that seems inexpensive in testing can become the largest line in an infrastructure bill at production traffic. Model this before you launch, not after.
Fluency is not accuracy. This is the one to carry with you. These systems are extraordinarily good at sounding authoritative. That quality is completely independent of whether they are right, and it makes them uniquely easy to over-trust.