Degree 2 · Unit 2.5
Beyond text: image and sound
Once meaning had become numbers, the barrier between different media fell away. An image is turned into vectors too, and so is sound. Train a system on pairs of an image and its description, and it learns to place the image and the description at the same point in the space of meaning. After that it can describe an image it has never seen, or create an image from a description nobody ever wrote. This is multimodality, and there is nothing new in the principle behind it — what is new is how far it now reaches.
| Medium | Masters | Falls at |
| Image | Description, extracting text, reading tables | Exact counting, spatial relations, decisive details |
| Sound | Transcription, translation, reading tone | Local dialects, noise, overlapping speakers |
| Video | Summarising, searching within scenes | Long causal sequences, what is not said outright |
| Document | Reading, summarising, comparing | Poor scans, handwriting, badly encoded Arabic |
The other face: generation is forgery
The system that creates an image of a product that has not been manufactured yet is the very same system that creates an image of an event that never happened. And the system that reads your book aloud in your own voice, so that a partially sighted reader can follow it, is the very same system that produced the audio clip in the story that opened this programme. There are not two technologies here, one for good and one for harm. There is one technology and there are two intentions.
Machine creativity is born of an internal contest: one side forges, the other detects, and each raises the other.
fawzooz.ai
The loop runs thousands of times until the expert critic can no longer tell — and only then has the model "learned" to generate. The improvement is not a designer's gift but the product of mutual pressure; this is precisely the region where deepfake capability is born too.
This is why learning to use it is not enough; you must climb to the fourth degree. Mastering generation without understanding forgery is like learning to drive without ever being shown where the brakes are.
Do this
1 — On the tool. Photograph a page of a printed Arabic document and ask for it to be transcribed, then photograph a handwritten page. Compare the two results — the difference is the limit of the tool.
2 — On the tool. Give the system a photograph of a table from a real report and ask it to extract the figures. Check every figure by hand and record the error rate.
3 — In your field. If your voice or your face were used in a forgery tomorrow, what is the most dangerous thing that could be asked in your name? And who would believe it? Keep the answer for the fourth degree.
How to choose the right medium
A great deal of everyday failure comes from choosing the wrong medium for the task in hand. Photograph a table to have it analysed and you lose precision you would have kept simply by copying the figures out as text. Send a long recording to get a summary and you lose the details a written minute would have preserved for you.
The rule is straightforward: give the system the cleanest form of the information available to you, rather than the form that is easiest for you to hand over. Text is more precise than an image, a clear image is more precise than a poor scan, and a written minute is more precise than a recording. A minute spent improving the input will save you ten correcting the output.
And where the weak medium is all you have — a poor photograph, or handwriting — do not ask for a conclusion built on top of it. Ask for a transcription first, correct that transcription with your own eyes, and only then build on it. Separating extraction from inference is what stops an error from slipping quietly into your result.
Where to after this unit? One unit remains in this degree, and it is your bridge to what comes after: the three limits of the machine that no upgrade repairs.
