AI coursesMain page
Back to the library

Token embeddings and contextual representations

Connect discrete tokens to vectors and contextual computation.

On this page
  1. Token identifiers and embedding lookup
  2. Lookup arithmetic and gradients
  3. Position and contextual mixing
  4. Causal masking and shifted targets
  5. Relation to a complete language model
  6. References

Token IDs provide a discrete representation of text, but a network needs coordinates on which its learned transformations can operate. Embeddings provide those coordinates. Contextual computation then allows the representation of a token to change with the tokens around it, turning a reusable lookup vector into a representation suited to the current sequence.

Language models operate on numerical representations of tokens. Tokenization defines discrete units and their identifiers; embedding lookup assigns vectors to those identifiers; subsequent layers produce context-dependent representations. These stages have different mathematical roles and should not be conflated.

1. Token identifiers and embedding lookup

A tokenizer maps a string to a sequence of token IDs. Tokens may be words, subword units, punctuation or byte-derived units, depending on the tokenizer. An ID is an index, not a semantic measurement: numerical adjacency between IDs does not imply similarity between tokens. [1]

Let VV denote vocabulary size and dd embedding width. An embedding matrix E∈RV×dE\in\mathbb R^{V\times d} stores one vector per token. With zero-based ID ii and one-hot column vector ei∈RVe_i\in\mathbb R^V, lookup is:

xi=E⊤ei∈Rdx_i=E^\top e_i\in\mathbb R^d

The one-hot form expresses row selection algebraically. An implementation can index the row directly without allocating a length-VV vector. The embedding table may be trained or held fixed, depending on the specified procedure. [2]

2. Lookup arithmetic and gradients

Assign IDs zero, one and two to abstract tokens α,β,γ\alpha,\beta,\gamma, and define:

E=(101011100101)E=\begin{pmatrix}1&0&1&0\\1&1&1&0\\0&1&0&1\end{pmatrix}

For the sequence (2,0)(2,0), lookup gives:

X=(01011010)∈R2×4X=\begin{pmatrix}0&1&0&1\\1&0&1&0\end{pmatrix}\in\mathbb R^{2\times4}

There are 12 embedding parameters and eight entries in this sequence representation. The squared distances in this constructed table are ∥E0−E1∥22=1\|E_0-E_1\|_2^2=1 and ∥E0−E2∥22=4\|E_0-E_2\|_2^2=4, demonstrating how vector coordinates permit comparisons between token representations.

For a general length-TT sequence, write X=SEX=SE, where S∈{0,1}T×VS\in\{0,1\}^{T\times V} has one one-hot row per position. For scalar loss L\mathcal L and upstream gradient G=∂L/∂XG=\partial \mathcal L/\partial X, the chain rule gives:

∂L∂E=S⊤G\frac{\partial \mathcal L}{\partial E}=S^\top G

Repeated occurrences of the same ID accumulate gradient contributions in the same row of EE. This derivative is distinct from an optimizer update, which may include additional terms or leave the table frozen.

3. Position and contextual mixing

A fixed embedding table returns the same vector for repeated uses of an ID. Dependence on surrounding tokens arises from subsequent computation. A single self-attention head with additive positional representations P∈RT×dP\in\mathbb R^{T\times d} can be defined by:

H=X+P,Q=HWQ,K=HWK,Vh=HWVH=X+P,\quad Q=H\mathbf{W}_Q,\quad K=H\mathbf{W}_K,\quad V_h=H\mathbf{W}_V

The projection matrices have shapes:

WQ,WK∈Rd×dk,WV∈Rd×dv\mathbf{W}_Q,\mathbf{W}_K\in\mathbb R^{d\times d_k},\quad \mathbf{W}_V\in\mathbb R^{d\times d_v}

Thus Q,K∈RT×dkQ,K\in\mathbb R^{T\times d_k} and Vh∈RT×dvV_h\in\mathbb R^{T\times d_v}. The subscript distinguishes the value matrix VhV_h from vocabulary size VV. Each query is compared with the keys by dot products. Rowwise softmax converts these scores into nonnegative weights summing to one; those weights combine the value vectors. The factor 1/dk1/\sqrt{d_k} moderates the score scale as key width grows. Attention weights and outputs are:

A=softmax⁡row(QK⊤dk+M),O=AVhA=\operatorname{softmax}_{\mathrm{row}}\left(\frac{QK^\top}{\sqrt{d_k}}+M\right),\qquad O=AV_h

Softmax is applied separately to each row, giving A∈RT×TA\in\mathbb R^{T\times T} and O∈RT×dvO\in\mathbb R^{T\times d_v}. Each output row is a weighted sum of permitted value rows. This example uses additive position vectors to supply the order information needed by the attention computation. [3]

4. Causal masking and shifted targets

For autoregressive computation, position ii must not use a later position jj. This restriction is expressed by:

Mij={0,j≤i,−∞,j>iM_{ij}=\begin{cases}0,&j\le i,\\-\infty,&j>i \end{cases}

Since exp⁡(−∞)=0\exp(-\infty)=0, future positions receive zero probability after normalization. Each row retains at least its current position. After the remaining decoder computations, the representation at position ii predicts xi+1x_{i+1}. The target is therefore shifted relative to the input, consistent with p(xt∣x<t)p(x_t\mid x_{<t}) when t=i+1t=i+1.

For a two-position arithmetic example, let every unmasked score be zero and the value rows be (2,0)(2,0) and (0,4)(0,4). Then:

A=(101/21/2),O=(2012)A=\begin{pmatrix}1&0\\1/2&1/2\end{pmatrix},\qquad O=\begin{pmatrix}2&0\\1&2\end{pmatrix}

The first output uses only the first value; the second averages both because its two permitted scores are equal. The mask determines which positions can contribute, and the scores determine their relative weights.

5. Relation to a complete language model

Embedding lookup connects discrete IDs to vectors; contextual layers combine information across allowed positions; the output projection and softmax described in article 2 produce next-token probabilities. A static embedding and a contextual hidden state are therefore different objects.

A complete Transformer additionally specifies multiple heads, output projections, residual paths, normalization, feed-forward layers and their ordering. The original encoder–decoder Transformer is also distinct from a causal decoder language model. The single-head computation above supplies the attention component of that larger structure. [3]

Neither individual coordinates nor attention weights automatically provide complete semantic or causal explanations. Their interpretation must account for the surrounding transformations and the task for which the parameters were fitted.

References