Course Content
Introduction to AI
4 sections · 10 lessons
History & Evolution of AI
In 1970, one of the most respected AI researchers alive told Life magazine that within three to eight years we would have a machine with the general intelligence of an average human being.
He was wrong by at least half a century. And he was not a fool — he was brilliant, and he had good reasons for believing it.
Understanding why that prediction failed, and why similar predictions failed again in the 1980s, is genuinely useful to you. Not as trivia. The field has been through two full cycles of enormous excitement followed by collapse, and both collapses happened for reasons that are still relevant. Learning to spot the pattern is one of the more practical things this course can give you.
This lesson walks through seven decades. For each era, we will look at what people believed, what they actually built, and what stopped them.
Before the beginning: the question that started it
In 1950, the British mathematician Alan Turing published a paper that opened with a deceptively simple question: can machines think?
Turing immediately pointed out the problem with his own question. "Think" is too vague to test. So he replaced it with something you could actually run.
Put a person at a keyboard. Let them exchange typed messages with two hidden participants — one human, one machine. If the person cannot reliably tell which is which, then arguing about whether the machine "really" thinks is not a productive use of anyone's time.
This became known as the Turing Test, and its importance is not that it is a good measure of intelligence — most researchers now agree it is not. Its importance is the move Turing made: replacing an unanswerable philosophical question with a concrete, observable behaviour.
That instinct — stop asking whether a machine understands, and start asking what it can reliably do — remains the working philosophy of the entire field. It is also why AI progress is measured in benchmarks rather than in arguments.
1956–1974: the golden age
The summer that named the field
In the summer of 1956, a small group of researchers gathered at Dartmouth College for a workshop. Their proposal contained a sentence of extraordinary confidence:
"Every aspect of learning or any other feature of intelligence can in principle be so precisely described that a machine can be made to simulate it."
The term artificial intelligence was coined at that meeting. They estimated the core problems could be substantially solved over a summer by about ten people.
What they actually achieved
It is easy to laugh at that estimate now. It is fairer to notice how much they really did accomplish, with computers vastly less capable than a modern washing machine.
- Programs that proved mathematical theorems — one found a proof shorter and more elegant than the one in Principia Mathematica, the standard reference work of its day.
- Programs that played checkers well enough to beat their own creators, and improved with practice.
- ELIZA (1966), a program that imitated a psychotherapist by reflecting your statements back as questions. It was a few hundred lines of pattern matching with no understanding whatsoever — and people became genuinely attached to it, confided in it, and asked to be left alone with it. Its creator was disturbed by this for the rest of his life.
- Early robotics — machines that could see simple blocks and stack them on command.
Real results. So why did it stall?
The wall they hit
Three problems, and each one is still worth understanding.
Combinatorial explosion. These programs worked by searching possible sequences of moves or steps. In a small toy world, that is fine. But the number of possibilities grows explosively with problem size. Chess has roughly 10120 possible games — more than there are atoms in the observable universe. You cannot search your way through that, and no faster computer rescues you, because the numbers grow faster than hardware ever will.
Common sense turned out to be enormous. Researchers assumed the hard part was logic and the easy part was ordinary knowledge. It was exactly backwards. Proving a theorem is comparatively easy to encode. Knowing that water is wet, that a person who leaves a room is still alive, that you cannot push a rope — this is a vast, mostly unwritten body of knowledge that humans acquire by living.
Toy problems did not scale. A program that stacked coloured blocks in a simulated world could not be extended to a real table with real lighting and real clutter. The simplified world was not a small version of reality; it was a different thing.
The first AI winter (1974–1980)
By the early 1970s, funders had noticed the gap between promises and delivery. Reports were commissioned. They were damning. Government funding in Britain and the United States was cut sharply. Researchers left the field or renamed their work to avoid the term "AI" on a grant application.
This period is called the first AI winter, and the mechanism is worth naming precisely, because it repeats:
Early success on a narrow problem → confident extrapolation to the general problem → funding based on the extrapolation → the general problem turns out to be qualitatively harder → collapse of confidence.
1980–1987: expert systems and the second boom
A new, more modest idea
The next wave gave up on general intelligence and asked a narrower question: instead of building a mind, could we capture what one expert knows about one subject?
The result was the expert system: a large collection of hand-written rules extracted from human specialists through lengthy interviews.
IF the organism is gram-positiveAND the organism grows in chainsAND the infection site is bloodTHEN there is suggestive evidence (0.7) that the organism is streptococcusStack thousands of rules like this and you get a system that can work through a genuine diagnostic problem. And these systems worked. One medical system outperformed junior doctors at diagnosing blood infections. A configuration system at Digital Equipment Corporation reportedly saved the company tens of millions of dollars a year.
Money poured in. A whole industry of specialised hardware grew up to run these systems.
Why it collapsed anyway
Expert systems failed for reasons that had nothing to do with being wrong, and everything to do with being unmaintainable.
| Problem | What it meant in practice |
|---|---|
| Knowledge acquisition bottleneck | Every rule required interviewing an expert. Experts are expensive, busy, and often cannot articulate what they know. |
| Brittleness | Given a case slightly outside its rules, the system did not degrade gracefully — it produced confident nonsense. |
| Maintenance cost | Medicine changes. Every change meant re-interviewing experts and checking thousands of interacting rules. |
| No learning | The system never improved from experience. It knew exactly what it was told, forever. |
By 1987 the specialised hardware market collapsed — cheaper general-purpose workstations had caught up. The second AI winter followed, lasting into the mid-1990s.
The lesson that survived: knowledge that has to be entered by hand does not scale, and a system that cannot learn cannot keep up with a changing world. This realisation set up everything that came next.
1990s–2010: the quiet, decisive shift
The winter was not empty. Underneath the disappointment, the field changed its fundamental approach — and this shift matters more than any single achievement in this lesson.
Researchers stopped trying to encode knowledge and started trying to learn it from data. Statistics and probability replaced hand-written logic. The field of machine learning moved from the margins to the centre.
The change is easiest to see in translation. The old approach hand-coded grammar rules for each language pair — years of linguist time, mediocre results. The new approach fed the system millions of documents already translated by humans and let it learn which phrases correspond to which. It knew no grammar at all. It worked far better.
What made the shift possible
Three things arrived at once, and none of them was a clever new idea:
- Data. The internet made enormous text and image collections available for the first time in history.
- Computing power. Processors kept getting faster and cheaper, year after year.
- Storage. Keeping millions of examples became affordable rather than exotic.
Notice what is not on that list: fundamentally new theory. Many of the key algorithms had been published decades earlier and simply could not be run at useful scale.
Deep Blue, 1997
In 1997, IBM's Deep Blue beat world chess champion Garry Kasparov. Headlines declared machines had surpassed human intelligence.
They had not, and the details matter. Deep Blue searched around 200 million positions per second using evaluation rules written by human grandmasters. It was a magnificent piece of engineering and a very fast, very narrow search. It could not play draughts, hold a conversation, or recognise a chessboard in a photograph.
The real lesson of Deep Blue is one people still get wrong: superhuman performance on a narrow task tells you very little about general capability.
2012–2017: deep learning arrives
The result that changed everything
In 2012, a neural network called AlexNet entered an annual image recognition competition. Previous winners had error rates around 26%, improving by a percentage point or so each year.
AlexNet scored 15.3%. It did not edge ahead — it collapsed the error rate by roughly ten percentage points in a single year.
The field noticed immediately. Within three years, essentially every serious entry was a neural network.
Why then, and not twenty years earlier?
Here is the genuinely surprising part: the core ideas behind AlexNet were decades old. Neural networks date to the 1950s. The training algorithm that makes them work was popularised in 1986.
Three things had changed, and none of them was the theory:
| Ingredient | What changed | Why it mattered |
|---|---|---|
| Data | ImageNet: 14 million hand-labelled images | Deep networks need enormous numbers of examples |
| Hardware | Graphics cards repurposed for training | Cut training from months to days |
| Technique | Refinements that stopped deep networks failing to train | Made depth practical rather than theoretical |
Graphics cards deserve a moment of appreciation. They were built to render video game visuals, which requires doing the same simple arithmetic on millions of pixels simultaneously. Training a neural network requires doing the same simple arithmetic on millions of numbers simultaneously. The hardware built for entertainment turned out to be almost perfectly shaped for AI — an accident that shaped the next decade.
What followed quickly
- 2014 — networks that generate realistic images by pitting two networks against each other, one producing fakes and one detecting them.
- 2016 — AlphaGo beat a world champion at Go, a game with vastly more possible positions than chess. Unlike Deep Blue, it could not brute-force search; it learned intuition about which positions were promising.
- 2017 — a paper titled Attention Is All You Need introduced the Transformer architecture. Nearly every major AI system you have heard of since is built on it.
2018–today: the generative era
The Transformer made it practical to train on truly enormous amounts of text. Researchers then found something unexpected.
Make the model bigger, give it more data and more computing power, and capabilities appear that nobody explicitly trained for. A model trained only to predict the next word in a sentence turns out to be able to summarise, translate, answer questions, and write working code — none of which were training objectives.
This is genuinely strange, and it is worth sitting with. The task is mechanical: given this text, what word comes next? But doing that task extremely well across a large fraction of everything humans have written apparently requires internalising a great deal about how the world is described.
The moment it reached everyone
Powerful language models existed for several years before most people encountered one. What changed in late 2022 was not primarily capability — it was interface. Putting a model behind a chat box that anyone could type into took AI from a research topic to something, by widely cited estimates, a hundred million people used within two months.
Worth remembering as you build things: the technology had existed for years. The thing that changed the world was making it usable. Access and interface are not lesser problems than modelling — they are frequently the bottleneck.
After the chat box: seeing, reasoning, acting
The years since have brought three extensions of the same idea rather than a new one. You will meet all three constantly, so it is worth knowing them by name.
Models that take in more than text. Leading models are now multimodal: they accept photos, screenshots, documents and speech alongside typed text, and many can reply in speech or produce images. The approach has not changed — learn patterns from enormous numbers of examples — it is simply applied to several kinds of data at once.
Models that work through a problem before answering. From late 2024, a new kind of model, usually called a reasoning model, was trained to write out a long chain of intermediate working before giving its final answer, and rewarded when that working led to a correct result. OpenAI's o1 (2024) and DeepSeek's R1 (early 2025) were early public examples, and most leading models now offer a mode like this. Spending more computation at answer time brought large gains on maths, science and coding problems. It reduced mistakes; it did not remove them.
Models that take actions. An agent is a model placed in a loop with tools — web search, running code, editing files, operating a browser — so it can carry out a multi-step task instead of only replying. Coding agents that read a codebase, make changes and run the tests became widely used during 2025. An agent inherits every weakness of the model inside it and adds a new one: a wrong step now does something, rather than just saying something.
The pattern across seven decades
Step back and the same shape appears three times.
| Era | The big idea | What stopped it |
|---|---|---|
| 1956–1974 | Encode reasoning as search | Possibilities explode; common sense is vast |
| 1980–1987 | Encode expert knowledge as rules | Rules do not scale and cannot learn |
| 1990–2010 | Learn patterns from data | Limited by available data and computing power |
| 2012–2017 | Deep networks at scale | Hungry for data, expensive, hard to explain |
| 2018–now | Very large generative models | Cost, reliability, and confident invention |
Three durable lessons come out of this history.
Progress comes from resources as often as from ideas. The 2012 breakthrough used decades-old theory with new data and new hardware. When you hear about a breakthrough, ask what became newly possible, not only what became newly thought.
Narrow success is a poor predictor of general capability. Chess in 1997, Go in 2016, image labelling in 2012 — each was read as a sign of imminent general intelligence. Each was a narrow system that could not do anything else.
Winters follow overpromising, not underperformance. Both collapses came after real technical achievements. What broke was the gap between what was promised and what was delivered. This is a useful lens for evaluating claims you will encounter throughout your career.