Course Content
Introduction to AI
4 sections · 10 lessons
AI in Healthcare, Finance, Education, and Entertainment
It is easy to read about AI in healthcare or finance and come away with a vague impression that "computers help doctors now". That impression is nearly useless. It tells you nothing about what the software actually does, where it fails, or why some deployments succeed while others quietly get switched off.
This lesson takes four industries and looks underneath. For each one we will ask the same four questions: what specific task is being automated, what does the system see, where does it break, and what happened when someone tried it for real?
The goal is not to memorise applications. It is to develop an eye for the difference between AI that works and AI that merely demonstrates well.
Healthcare
Reading medical images
This is the most mature use of AI in medicine, and the reason is structural: it is a supervised learning problem with an unusually clean shape. The input is an image. The output is a label. Hospitals have decades of scans, each already reviewed by a specialist who wrote down what they found.
Systems in real use today screen for signs of diabetic eye disease, flag suspicious areas on mammograms, spot lung nodules on chest scans, and triage brain scans so the ones showing a possible bleed jump the queue.
That last one is worth pausing on, because it shows a design choice that separates useful medical AI from the demo kind.
A stroke triage system does not diagnose. It reorders the queue. A radiologist still reads every scan and makes every call — but the scans most likely to be urgent arrive at the top of their list instead of the middle. Nobody is replaced, no diagnosis is automated, and minutes are saved where minutes cause permanent disability.
The system is allowed to be wrong. If it promotes a scan that turns out to be fine, a radiologist looks at a normal scan slightly earlier than they would have. The cost of a false alarm is close to zero. That asymmetry is what makes the deployment safe.
Where medical AI actually fails
The failures are rarely "the model was inaccurate in testing". They are almost always about the gap between the test and the world.
It learned the hospital, not the disease. A now-famous class of failure: a model trained to detect pneumonia performed superbly in testing and collapsed at a new hospital. Investigation found it had partly learned to recognise the scanner. Sicker patients had been imaged on a particular portable machine, which left a subtle signature on the image. The model found that signature and used it. It was never looking at lungs the way anyone assumed.
The population shifted. A model trained mostly on one demographic performs worse on others. Skin lesion classifiers trained overwhelmingly on light skin have measurably worse accuracy on dark skin. This is not a bug in the algorithm; it is the data being unrepresentative, faithfully reproduced.
Nobody used it. The most common failure of all, and the least written about. A model that produces an alert clinicians do not trust, in a system they have to leave their workflow to check, gets ignored within a fortnight. Accuracy is necessary and nowhere near sufficient.
What AI does not do in medicine
- Decide treatment. It can surface options and flag risks; a clinician decides and is accountable.
- Take a history. Much of diagnosis comes from a conversation, including what the patient is reluctant to say.
- Weigh a patient's life. Whether an eighty-year-old should undergo aggressive treatment is not a prediction problem.
- Carry responsibility. If an automated decision harms someone, "the model said so" is not a defence.
Finance
Finance adopted machine learning earlier than almost any other industry, for an unromantic reason: the data was already digital, already labelled, and directly attached to money.
Fraud detection
Every card transaction is scored in the time it takes the terminal to respond — typically under a hundred milliseconds. The model considers the amount, the merchant, the location, the time, the device, and crucially how all of that compares to your own history.
Fraud detection is a useful teaching example because of a property that makes it genuinely hard:
Extreme class imbalance. Perhaps one transaction in two thousand is fraudulent. A model that simply answers "not fraud" every single time is 99.95% accurate — and completely worthless. This is the clearest demonstration you will find that accuracy alone is a misleading measure.
What matters instead is the trade between two kinds of error, and the trade is not symmetric:
| Error | What happens | Cost |
|---|---|---|
| Missed fraud | A genuine fraud goes through | The bank refunds the money |
| False alarm | Your real payment is declined | Embarrassment, a support call, sometimes a lost customer |
Banks tune this deliberately. Too sensitive and customers are declined at supermarket tills and leave. Too permissive and losses climb. The dial is a business decision informed by the model, not a technical one the model makes.
Credit scoring, and why it is regulated
Credit models predict the probability that a borrower defaults. They work well. They are also the area where AI regulation bites hardest, and the reason is instructive.
A model may not legally use race in a lending decision. But a model given postcode, shopping patterns and employment history can reconstruct a close approximation of race without ever being told it — because those things correlate in a society with a history of segregation. The technical term is a proxy variable, and it is the single most important idea in fairness work.
Removing a protected attribute from the input does not remove it from the model. The information leaks in through correlated variables. This is why "we don't collect race, so we can't be biased" is one of the most confidently wrong statements in the industry.
Regulation in many jurisdictions requires that a rejected applicant can be told why. That requirement rules out the least explainable models regardless of their accuracy, and it is a good example of a constraint that is not technical at all shaping which technique you may use.
Algorithmic trading
Models predict short-term price movements and execute automatically. Worth knowing for one cautionary reason: financial markets are the clearest example of an environment that reacts to your model. A pattern that reliably predicted returns stops working once enough participants trade on it — the profit is what destroys the signal. Most machine learning assumes the world does not adapt to your predictions. Here it does, quickly.
Education
What genuinely works
Adaptive practice. The best-evidenced use of AI in education. The system tracks which problems you get right, estimates what you have and have not yet mastered, and chooses the next problem to sit just past your current ability. Get it wrong and it steps back to a prerequisite; get it right repeatedly and it moves on.
The underlying method predates deep learning by decades and does not need a large model. It works because it solves a real constraint: one teacher cannot give thirty students individually paced practice, and software can.
Fast, low-stakes feedback. Automatic marking of code, maths, and structured answers means a student finds out they misunderstood something in seconds instead of a week. Speed of feedback matters enormously for learning, and it is a pure throughput problem — exactly what software is for.
Access. Real-time captioning, translation, text-to-speech and reading support let students participate who otherwise could not. This is unglamorous and probably the highest-value application in the whole list.
What is oversold
Essay grading. Automated scorers correlate reasonably with human markers on average — but they are measuring surface features like length, vocabulary range and structure. Feed them fluent, well-structured nonsense and they score it highly. As feedback on mechanics they are useful. As a judgement of whether a student understood something, they are not.
"AI tutors" that answer everything. A model that instantly supplies a correct-looking answer can prevent the productive struggle that actually produces learning. There is also the confident-falsehood problem: a student without the expertise to spot a wrong answer is precisely the person least equipped to catch it.
Attention monitoring. Systems claiming to detect engagement from a webcam rest on shaky science and raise serious surveillance concerns. Treat these claims with suspicion.
Entertainment
Recommendation, and what it optimises
Recommendation is the most commercially consequential AI most people interact with daily. A large fraction of what gets watched on major streaming platforms comes from a recommendation rather than a search.
The mechanics combine two ideas. Collaborative filtering finds people whose behaviour resembles yours and suggests what they liked. Content-based filtering describes each item by its properties and matches those to your history. Real systems blend both, plus recency, popularity, and business rules.
The part worth thinking hard about is not the technique but the objective:
A recommender optimises whatever you tell it to. Optimise for watch time and you will get watch time — including through content that is compelling in ways nobody would defend. The model is not misbehaving; it is succeeding at the goal it was given. Choosing the objective is where the ethics live, and it is a human decision.
Generation in creative work
In practice, generative tools have landed mostly in the unglamorous middle of production pipelines: filling in background elements, generating variations for a client to choose between, cleaning up audio, upscaling old footage, drafting rough concepts before a human refines them.
Two genuine controversies you should be able to discuss rather than dismiss.
Training data consent. Image and text models were largely trained on material scraped from the internet, much of it copyrighted, without the creator's permission or payment. Lawsuits have followed in several countries; some have produced early rulings or settlements, but courts have not converged and the broad question is still unsettled.
Synthetic likeness. Realistic voice and video generation makes convincing fabrication cheap. The technical countermeasures — watermarking, detection models — are in a permanent arms race with the generators, and detection is losing more often than it wins.
The pattern underneath all four
Look across the industries and the same shape appears in every successful deployment.
| Property | Why it decides success |
|---|---|
| Narrow, well-defined task | "Flag urgent scans" succeeds where "assist doctors" has no measurable target |
| Plenty of labelled history | Past decisions supply the training answers for free |
| Cheap errors, or a human catching them | Nothing irreversible happens on a single prediction |
| Volume beyond human capacity | Millions of transactions per hour is the actual problem being solved |
| It fits the existing workflow | An alert nobody sees changes nothing, however accurate |
And the inverse — the conditions under which AI projects fail, which are remarkably consistent:
- The task was never defined precisely enough to measure.
- The training data did not resemble the situation where it was deployed.
- A single wrong prediction could do irreversible harm, with no human in the loop.
- The people expected to use it were not involved in designing it.
- Accuracy was the only thing measured, and accuracy was the wrong measure.
How to evaluate any AI claim
You will encounter claims like "our AI diagnoses skin cancer better than dermatologists" throughout your career. Six questions cut through nearly all of them.
- What exactly is the task? "Better than dermatologists" at what — one condition, from one kind of photograph, on which patients?
- Better by what measure? Accuracy hides the imbalance problem. Ask about both kinds of error separately.
- Tested on whom? Same hospital as training, or somewhere genuinely different? Which demographics?
- What happens when it is wrong? Who catches it, and how quickly?
- Compared against what? Against a specialist, or against the status quo of nothing at all? Those are very different claims.
- Is it deployed? Working in a live hospital for two years and winning a benchmark are not the same achievement.