You search for "holiday home" in the company's document archive and get nothing, even though there are ten documents about summer cottages in there. Traditional search compares characters, not meaning. Semantic search solves this with embeddings: text is translated into numbers so that sentences with similar meaning end up close to each other, regardless of which words were used. For a company, that means everything that has already been written, manuals, quotes, campaigns and replies, finally becomes findable.
Embeddings: text becomes a point in space
An embedding is a long sequence of numbers, a vector, that a language model computes from a piece of text. Think of it as a set of coordinates: every sentence, paragraph or document gets a point in a space with hundreds of dimensions. The model is trained so that texts that mean similar things get points that sit close together, while unrelated texts end up far apart.
That is why "summer cottage" and "holiday home" are neighbours in this space even though the words themselves have almost nothing in common. The model has seen them used in the same context, about the same thing, and so places them in the same region. A search that compares characters can never see this connection. A search that compares positions in meaning space sees it immediately.
The same applies to whole sentences. "How do I cancel my subscription?" and "I want to stop using the service" have almost no words in common but end up close to each other, because the meaning is the same. This is what makes semantic search something other than a fancier version of keyword search: it works with what you mean, not what you typed.
How is closeness measured?
Once the texts have become vectors, you need a number that says how similar two of them are. The most common measure is cosine similarity:
$$\text{similarity}(a, b) = \frac{a \cdot b}{\lVert a \rVert , \lVert b \rVert}$$
The formula looks technical, but the idea is simple. Think of the two vectors as arrows from the same starting point. Cosine similarity measures the angle between them: arrows pointing in the same direction get a similarity close to 1, which means the same meaning, while arrows pointing in different directions get a low number. The numerator, the dot product $a \cdot b$, measures how much the arrows work together, and the denominator corrects for their length so that only the direction matters. A long report and a short sentence on the same subject can therefore measure as very similar.
The search itself is then simple at its core: the user's question is turned into a vector, that vector is compared with the vectors of every document in the collection, and the documents that measure as most similar are brought forward.
What can companies do with this?
The most valuable dataset most companies have is the material their staff have already written: past marketing campaigns, product copy, procedures, manuals and thousands of replies to customer enquiries. The problem is not that the material is missing, but that nobody can find it. Semantic search changes that, and the same underlying technology opens more doors:
- Search in your own material. The marketing team finds every past campaign on a given theme, a customer service representative finds the right answer even when the customer phrases the question differently from the manual.
- Duplicate clean-up. Two documents that say almost the same thing measure as highly similar even when the wording differs. That makes it possible to find duplicated content in a document archive or product catalogue and keep a single correct version.
- Sorting customer feedback. Hundreds of open-ended survey responses are automatically grouped into clusters by meaning: everything about delivery times ends up together, everything about pricing together, without anyone defining the categories in advance.
- An assistant that answers from your own data. Before a language model answers a question, semantic search is used to retrieve the relevant internal documents, and the model then answers based on them instead of guessing. This pattern, retrieve first and answer second, is called retrieval-augmented generation (RAG) and is the most common way of getting AI to answer correctly about material it was never trained on.
What to watch out for
The technology is powerful, but it is not magic. Three things matter most in practice:
- Embeddings are tied to the model that created them. Vectors from two different models cannot be compared, like coordinates from two different map systems. If you switch models, the entire collection has to be recomputed, and it pays to know that before the collection runs to millions of documents.
- Quality in Icelandic has to be tested. Many embedding models are trained overwhelmingly on English. Some handle Icelandic well, others do not, and the only way to know is to test on your own data: take real queries and check whether the right documents come out on top.
- The quality of the search determines the quality of the answer. In a RAG setup, the model answers based on what the search returns. If the search brings up the wrong or an outdated document, the model delivers a confident answer built on the wrong foundation. Garbage in, convincing garbage out. That is why the search component, not the language model, is usually the place where care makes the biggest difference.
What you can do right now
- Write down five real searches that failed. Ask the team when they could not find a document, an answer or older material they knew existed. This list is both the business case and the test set for semantic search.
- Map where your text lives. Manuals, quotes, campaigns, replies to enquiries: where is the material stored and in what format? Semantic search is only as good as its access to the material it is supposed to search.
- Test on a small collection before you buy a big system. Take one well-defined collection, for example the most common customer service replies, and have someone test whether semantic search finds the right answer to real enquiries in Icelandic. The result will tell you more than any sales pitch.
[ Get in touch ]
Book a free assessment
90 minutes that pay off immediately: we map your AI usage, risks and 3 to 5 automatable workflows, and deliver a report within a week. No commitment.