Degree 2 · Unit 2.2
How a machine learns
Inside every system that learns — from the spam filter on your email to the largest language model in the world — the same loop is running, millions of times over. It has four steps, and there is no fifth.
Repeat that loop billions of times across trillions of examples, and what you get is what we call artificial intelligence today. There is no magic in it at all, only a vast amount of computational patience.
Choosing how the model learns is the first strategic decision in any project — and the one that shapes the result most.
fawzooz.ai
The right question is not "which algorithm is stronger?" but "which form of supervision does my data actually allow?". And most of a project's time goes to cleaning data, not to building the model.
Data governs everything
If the loop is the same in every system, what is it that separates a useful system from a harmful one? The answer is what the machine was shown. Data is not neutral fuel; it is the teacher. Every flaw in the data becomes a flaw in the system, except that it comes out polished, in confident language that hides where it came from.
Three flaws come up again and again. The first is absence: a group that is missing from the data will be the group the system always gets wrong. The second is bias: if the decisions of the past leaned one way, the system learns to repeat that lean and then presents it as objectivity. The third is the false shortcut, where the machine finds an easy marker that lands on the right answer without any understanding behind it.
AI without data is an engine without fuel — and this is the fuel cycle, whole.
fawzooz.ai
The cycle is not a straight line: evaluation often exposes a flaw in the data and sends you back to the second stage. That is the process succeeding, not failing.
Memorising versus learning
A student who memorises last year's exam questions does brilliantly on those questions and fails at everything else. A machine does exactly the same thing, and we call it overfitting: it masters the training data perfectly and then fails outside it. This is why no system is ever measured on the data it learned from, but on data it has never seen before.
You do not build models, so there is one practical sign to watch for: a striking performance in the demonstration, and a modest one in your actual working day. So when a vendor offers you a system with impressive accuracy, ask them a single question — what data was that figure measured on, and where did that data come from?
A deep model does not see the image all at once; it builds it from the bottom up.
fawzooz.ai
No single layer "understands" anything. Understanding is a property of the arrangement, not of the parts — which is why a model's decision cannot be explained by taking one cell apart.
Do this
1 — On paper. Explain the four-step training loop to a member of your family, without one foreign term. If they do not understand, the fault is in your understanding, not in your explanation.
2 — On the tool. Ask an AI system to describe "the typical professional" in your field. Read the description critically: what bias did it inherit from its data? Which group is missing from its picture?
3 — In your field. Imagine a system learning from your organisation's decisions over the past five years. Write down three wrong patterns it would learn and repeat with confidence.
Where to after this unit? You now understand learning in general. The next unit goes into the kind you use every day: language models, and how predicting a single word manages to produce all of this.
