Language models: probability, training and generation
From next-token probabilities to a complete answer.
On this page
Generating a sentence requires choosing among possible continuations at each step. A probability distribution expresses how the model ranks those alternatives; training changes that distribution so observed continuations receive greater probability. This connects the next-token objective to the repeated computation that produces a complete answer.
An autoregressive language model defines conditional probabilities over discrete token sequences. Training estimates its parameters from observed sequences; generation selects successive tokens from the resulting distributions. These operations share a probability model but differ in both their inputs and their effects on the parameters.
1. Sequence probability
Let be the vocabulary, its size, and a sequence. The probability chain rule gives:
The first factor is conditioned on an empty prefix or an explicitly specified start token. For conditional generation, the supplied prompt forms part of the prefix. An end token can represent termination within the probability model. [1]
The factorization follows from conditional probability and permits dependence between positions. A Transformer can parameterize the conditional distributions. With a finite context limit , each computation uses the prefix admitted by that configuration.
2. From hidden states to probabilities
Suppose is the final hidden vector used to predict from its prefix. The output projection produces a vector of unnormalized scores, or logits:
Softmax converts these scores into probabilities:
For finite logits, each probability is positive and their sum is one. Direct exponentiation can overflow numerically. Subtracting the largest logit, , yields the equivalent expression:
The factor cancels between numerator and denominator. This transformation improves numerical evaluation without changing the represented distribution. [2]
3. Training objective and gradient
Maximum-likelihood training maximizes the probability assigned to observed sequences. Taking the negative logarithm converts a product into a sum. A token-average objective for one sequence is:
Averaging over tokens and averaging per-sequence losses produce different sample weights when sequence lengths vary. During next-token training, the observed prefix supplies the condition and the observed next token supplies the target. Causal attention excludes future positions from the computation used to predict a target. [3]
At a single position with target index , the negative log-likelihood expands to:
Differentiating the linear term and the log-sum-exp term gives:
Here equals one for the target class and zero otherwise. This is the derivative with respect to a logit. Parameter gradients additionally require the chain rule through the output projection and preceding layers.
4. A numerical example
For a constructed three-token vocabulary, let . Exponentiation gives , so normalization gives:
When the first token is the target, the loss is , and:
The calculation concerns token occurrence under a particular context. The first probability of therefore describes a possible continuation, not the factual correctness of a statement.
5. Generation and decoding
Decoding may select the most probable token, , or sample from the distribution. A positive temperature modifies sampling probabilities through . For the example above, changes the first probability to .
The selected token is appended to the prefix, and the computation repeats until a stopping condition is reached. During ordinary inference, remains fixed. A new prefix changes the conditional distribution without performing a training update. External storage, retrieval and subsequent training are separate system operations.
The diagram describes the computational order. The equations specify the distribution, training objective and distinction between inference and parameter updates. [3]