Token embeddings and contextual representations
Connect discrete tokens to vectors and contextual computation.
On this page
Token IDs provide a discrete representation of text, but a network needs coordinates on which its learned transformations can operate. Embeddings provide those coordinates. Contextual computation then allows the representation of a token to change with the tokens around it, turning a reusable lookup vector into a representation suited to the current sequence.
Language models operate on numerical representations of tokens. Tokenization defines discrete units and their identifiers; embedding lookup assigns vectors to those identifiers; subsequent layers produce context-dependent representations. These stages have different mathematical roles and should not be conflated.
1. Token identifiers and embedding lookup
A tokenizer maps a string to a sequence of token IDs. Tokens may be words, subword units, punctuation or byte-derived units, depending on the tokenizer. An ID is an index, not a semantic measurement: numerical adjacency between IDs does not imply similarity between tokens. [1]
Let denote vocabulary size and embedding width. An embedding matrix stores one vector per token. With zero-based ID and one-hot column vector , lookup is:
The one-hot form expresses row selection algebraically. An implementation can index the row directly without allocating a length- vector. The embedding table may be trained or held fixed, depending on the specified procedure. [2]
2. Lookup arithmetic and gradients
Assign IDs zero, one and two to abstract tokens , and define:
For the sequence , lookup gives:
There are 12 embedding parameters and eight entries in this sequence representation. The squared distances in this constructed table are and , demonstrating how vector coordinates permit comparisons between token representations.
For a general length- sequence, write , where has one one-hot row per position. For scalar loss and upstream gradient , the chain rule gives:
Repeated occurrences of the same ID accumulate gradient contributions in the same row of . This derivative is distinct from an optimizer update, which may include additional terms or leave the table frozen.
3. Position and contextual mixing
A fixed embedding table returns the same vector for repeated uses of an ID. Dependence on surrounding tokens arises from subsequent computation. A single self-attention head with additive positional representations can be defined by:
The projection matrices have shapes:
Thus and . The subscript distinguishes the value matrix from vocabulary size . Each query is compared with the keys by dot products. Rowwise softmax converts these scores into nonnegative weights summing to one; those weights combine the value vectors. The factor moderates the score scale as key width grows. Attention weights and outputs are:
Softmax is applied separately to each row, giving and . Each output row is a weighted sum of permitted value rows. This example uses additive position vectors to supply the order information needed by the attention computation. [3]
4. Causal masking and shifted targets
For autoregressive computation, position must not use a later position . This restriction is expressed by:
Since , future positions receive zero probability after normalization. Each row retains at least its current position. After the remaining decoder computations, the representation at position predicts . The target is therefore shifted relative to the input, consistent with when .
For a two-position arithmetic example, let every unmasked score be zero and the value rows be and . Then:
The first output uses only the first value; the second averages both because its two permitted scores are equal. The mask determines which positions can contribute, and the scores determine their relative weights.
5. Relation to a complete language model
Embedding lookup connects discrete IDs to vectors; contextual layers combine information across allowed positions; the output projection and softmax described in article 2 produce next-token probabilities. A static embedding and a contextual hidden state are therefore different objects.
A complete Transformer additionally specifies multiple heads, output projections, residual paths, normalization, feed-forward layers and their ordering. The original encoder–decoder Transformer is also distinct from a causal decoder language model. The single-head computation above supplies the attention component of that larger structure. [3]
Neither individual coordinates nor attention weights automatically provide complete semantic or causal explanations. Their interpretation must account for the surrounding transformations and the task for which the parameters were fitted.