Degree 2 · Unit 2.4
Tokens, vectors and context
A computer does not see letters at all. The first thing that happens to your text is that it gets cut into small pieces called tokens, and a token may be a whole word, part of a word, or a single punctuation mark. Each token is then turned into a list of numbers called a vector, which describes where that token sits in a mathematical space of meanings.
The elegant idea behind that space is that things close in meaning end up close in position. "Doctor" and "nurse" are neighbours; "doctor" and "cement" are a long way apart. This is how a system can understand the synonym you used even when you did not use its word, and how it notices a resemblance between two documents that do not share a single term. That representation is what makes semantic search possible, and you will build your knowledge assistant on it in the third degree.
The Arabic tax
Tokenisers were built mostly on English text, so a common English word became a single token while an Arabic word gets cut into three or four. The practical result is that Arabic text consumes roughly twice what the English equivalent does, both of the window and of the cost, and it loses more precision around diacritics and morphology.
You are not going to fix this on your own, but you can manage it. Shorten what you put in, clear the padding out of your documents, use models that state proper Arabic support, and do not be surprised when a system performs less well in Arabic than in English on the tasks that demand precision.
The context window: the memory of the moment
The context window is everything a system can see at any one moment: your instructions, your documents, the conversation up to that point, and the answer it is currently producing. It is not permanent memory. It is a workspace that gets wiped clean with every new conversation.
This window has widened enormously, until it can hold entire books, and a good many people assumed that settled the problem. It did not. The fuller the window gets, the weaker the system's attention becomes to whatever sits in the middle of it, and the more it favours the beginning and the end. That gives you a rule you will use constantly: put the most important instructions at the start or at the end, and never bury them in the middle.
Do this
1 — On paper. Write five words from your speciality and guess which will be cut into more tokens: the common ones or the rare ones? And why?
2 — On the tool. Give the system a long document with a small instruction buried in the middle of it. Did it obey? Repeat the experiment with the instruction at the end.
3 — In your field. Take a document you use often and cut it to half its size without losing a single piece of information.
A note for the Arabic reader
This is the unit where an Arabic reader realises that part of the weakness in their results comes from neither them nor their understanding, but from the structure of the tool itself. That should not become an excuse; it should become knowledge. Someone who knows their tool is heavier in their own language prepares for it accordingly: they put in cleaner text, they ask for a shorter output, and they verify twice where others verify once.
And this is precisely where an opportunity sits, one I will come back to in the sixth degree: the language gap that costs you today is the same gap that people who do not write in Arabic cannot even see.
Where to after this unit? You know how text is represented. The next unit leaves text for image and sound — and for the most dangerous thing in them.
