[ Blog ]

Tokens and embeddings: how language models read text

  • 11 September 2026
  • 5 min read

When you type a question to a language model (large language model, LLM), it does not read it the way a person does. It sees neither words nor sentences, but sequences of numbers. This sounds like a technical detail, but it is in fact the key to understanding how these tools behave: why they sometimes answer with complete confidence and are still wrong, why good instructions change everything, and why Icelandic costs more to use than English. A team that understands the machine uses it differently from a team that just bought access.

Step one: the text is chopped into tokens

The model's first job is to split the text into tokens. A token is not the same as a word. Common words often become a single token, while rarer words are broken down into smaller pieces: word stems, endings and even individual characters. The Icelandic word "fyrirtækjamenning" (company culture), for example, might become three or four tokens, while the English word "culture" will most likely be one.

Each token is then given a fixed ID number from the model's vocabulary. Your sentence has now become a sequence of numbers, for example [4021, 887, 15334, 92]. This is all the model gets to work with: no words, no grammar, just a series of numbers.

Step two: numbers that carry meaning

The ID numbers themselves say nothing about meaning. Number 4021 is no more "related" to number 4022 than to any other number. That is why each number is converted into an embedding: a long vector of numbers, often with hundreds or thousands of dimensions, which the model learned during training.

The remarkable thing about embeddings is that they capture meaning as distance. The vectors for "dog" and "cat" lie close to each other in this numerical space; "dog" and "interest rate" lie far apart. The model does not "know" what a dog is, but it knows that dog appears in similar contexts to cat in the enormous body of text it was trained on. Meaning, in the world of language models, is a position in numerical space.

One operation: predicting the next token

Here comes the part that surprises most people. Underneath all the features, the conversations, summaries and translations, a language model performs one and the same operation over and over: it predicts the next token based on the ones that came before. In the language of probability, it calculates

$$P(\text{next token} \mid \text{context})$$

for every single vocabulary number, arranges the results into a probability distribution (the last step in that calculation is called softmax, which turns raw numbers into probabilities that add up to 100%) and then picks a token from the distribution. That token is added to the context and the game repeats, token by token, until the answer is fully written.

This immediately explains two things. First: the model is not looking up facts, it is choosing a likely continuation. When the most likely continuation is correct, we get a clever answer. When the most likely continuation is wrong, we get convincing nonsense, because the sentence is just as well written in both cases. Second: everything in the context, the instructions, the examples, the data you paste in, changes the distribution $P(\text{next token} \mid \text{context})$ directly. Good instructions are not politeness towards the machine but a mathematical intervention in the prediction.

Icelandic is more expensive than English, quite literally

The tokenisation itself is learned, mostly from English text, because English dominates the training data of the large models. The result is that common English words get their own tokens while Icelandic words break down into more and smaller pieces. The same sentence therefore costs more tokens in Icelandic than in English.

This has two kinds of consequences for companies. On the one hand, price: most language models are priced by token count, so Icelandic use costs proportionally more for the same content. On the other hand, quality: a model that saw relatively little Icelandic during training has a weaker feel for inflections, word order and word choice than it has for English.

That is why you cannot assume that a model that is excellent in English is equally good in Icelandic. Performance has to be measured. The good news is that this is already being done: Miðeind maintains an Icelandic leaderboard where language models are compared on Icelandic tasks. For a company choosing tools for Icelandic-speaking staff, a measurement like that is a better basis than the vendors' marketing material.

What does this explain in everyday use?

Once tokens, embeddings and next-token prediction are clear, everyday phenomena in AI use become understandable:

  • Convincing nonsense is built in, not a malfunction. The model always picks a likely continuation, even when it has no reliable information. It does not say "I don't know" unless that sentence is the most likely continuation. That is why verifying facts, figures and references is a working rule, not optional caution.
  • The context is your control lever. The entire prediction rests on the context. Clear instructions, the right data and good examples shift the probability distribution towards the answer you want. A vague request returns the average of the internet.
  • Icelandic needs special attention. More tokens mean higher costs and often lower quality than in English. Choose tools based on measured performance in Icelandic, not on English demo videos.

This is the difference between a trained team and a tool buyer. The tool buyer is baffled when the model "lies". The trained team knows it was never telling the truth or lying, it was predicting tokens, and organises its way of working accordingly.

What you can do right now

  • Try a token counter. The makers of the large models offer free token counters online. Paste the same paragraph in Icelandic and English and compare the token counts. A five-minute exercise that makes pricing and context sizes tangible.
  • Build verification into the workflow, not into memory. Decide with the team which types of output always need to be verified: figures, names, legal references, prices. A written rule holds; "let's remember to check" does not.
  • Check Miðeind's Icelandic leaderboard before you choose a tool. If your staff work in Icelandic, performance in Icelandic is the prerequisite, not a side note. The measurement exists, use it.

[ Get in touch ]

Book a free assessment

90 minutes that pay off immediately: we map your AI usage, risks and 3 to 5 automatable workflows, and deliver a report within a week. No commitment.

No commitment

[ Direct contact ]

hallo@vestra.is+354 863 7496

Bolholt 8
105 Reykjavík, Iceland