[ Blog ]

Gradient descent: how AI learns

  • 11 September 2026
  • 5 min read

When a language model answers a question in flawless Icelandic, it is tempting to think someone programmed the grammar into it. Nobody did. The model learned it on its own, using one remarkably simple method repeated a remarkable number of times: guess, measure how wrong the guess was, and move a tiny step in the direction that reduces the error. This method is called gradient descent, and it is the heart of almost all modern AI. Anyone who understands it stops expecting magic and starts measuring.

Training means minimising error

A language model is, at its core, a gigantic formula with adjustable numbers. These numbers are called weights, and in the largest models there are hundreds of billions of them. Before training begins they are essentially random, and the model produces meaningless gibberish.

The training itself is a simple idea. The model is given text and has to guess the next word. The guess is then compared with the correct word, and the error is measured with what is called a loss function. The loss is just a number that says how wrong the guess was: a high number means a bad guess, a low number means a good guess. In words, one common loss function could be described like this: $L = $ how little probability the model gave to the correct word. If the model was fairly sure of the right answer, the loss is low; if the correct word caught it off guard, the loss is high.

All of training revolves around a single goal: finding the weights that make the loss as low as possible across an enormous collection of text. Nothing else. Grammar, facts and reasoning are never entered as rules; they are a by-product of minimising the error in next-word guesses.

Walking downhill in fog

The problem is that the search space is incomprehensibly large. With billions of weights there is no way to try every combination. This is where the classic analogy comes in: imagine you are standing on a mountainside in thick fog and want to get down to the valley. You cannot see the valley, not even ten metres ahead of you. The only thing you can feel is the slope under your feet.

The sensible approach is then obvious: find the direction in which the slope drops most steeply, take a small step that way and repeat. Step by step, without ever seeing the whole picture, you end up lower than where you started.

This is exactly what gradient descent does. The landscape is the loss function: every point in the landscape is one possible setting of all the weights, and the height at that point is the error that setting produces. The model never sees the whole landscape, but the mathematics gives it the slope where it stands, that is, the direction in which the error decreases fastest. It takes a small step in that direction and measures again. In the training of large language models this step is repeated an enormous number of times, on text running to billions of sentences.

The update rule, explained term by term

All of this work fits in a single line of mathematics, the update rule:

$$\theta_{t+1} = \theta_t - \eta , \nabla L(\theta_t)$$

It looks worse than it is. Let us take it apart:

  • $\theta_t$ is all of the model's settings, the weights, as they stand at step number $t$. Your current position on the mountainside.
  • $\nabla L(\theta_t)$ is the gradient of the error: the direction in which the loss grows fastest given the current settings. The minus sign in front reverses the direction, because we want to go down the slope, not up.
  • $\eta$ is the step size, usually called the learning rate. Too large a step and you leap over the valley and land on the slope on the other side; too small a step and the journey takes forever.
  • $\theta_{t+1}$ is then the new position: the old position plus one small step downhill.

Every single step nudges billions of weights ever so slightly, and the steps are taken across billions of sentences. No individual update teaches the model anything noticeable. But the sum of all these tiny nudges is a model that conjugates verbs correctly, translates between languages and writes summaries. The patterns settle into the weights without a single person having written a single rule.

What does this mean for your company?

This may sound like an academic curiosity, but three practical conclusions follow directly from how training works:

  • A new version is a new model, not an update of the old one. When a vendor releases a new version, behind it lies a new training run with new weights. A workflow that worked perfectly in the previous version can behave differently in the next one, for better or worse. That is why important workflows need to be retested whenever versions change.
  • A model's capability is measured, not promised. Nobody decided in advance what the model would know; the capability emerged as a by-product of error minimisation. The vendor itself finds out through testing what the model can handle. The same applies to you: the only way to know whether a model can handle your task is to test it on your data.
  • Fine-tuning and system prompts are different layers from the base training. Fine-tuning keeps nudging the weights with the same update rule, just on a smaller and more specific dataset. A system prompt changes no weights at all; it is text that steers behaviour in real time. Knowing which layer solves which problem saves both time and money.

Teams that understand this stop asking "why does nobody promise this will work?" and start asking "how do we measure whether this works for us?". That is a more mature and far more profitable question.

What you can do right now

  • Try the same task in two versions of the same model. Take a real task from your operations and run it in the latest version and an older one. The difference you see is tangible proof that new sets of weights behave differently.
  • Build a small test set for your most important workflow. Ten to twenty typical tasks with known correct answers are enough to start. Run the set at every version change, and you will base decisions on measurements rather than gut feeling.
  • Work out whether your problem calls for prompting or fine-tuning. If the model knows the task but does not follow your rules, a better system prompt is usually enough. Fine-tuning is only appropriate when the model lacks patterns that exist only in your data. Confusing the two is an expensive mistake.

[ Get in touch ]

Book a free assessment

90 minutes that pay off immediately: we map your AI usage, risks and 3 to 5 automatable workflows, and deliver a report within a week. No commitment.

No commitment

[ Direct contact ]

hallo@vestra.is+354 863 7496

Bolholt 8
105 Reykjavík, Iceland