Mathematical notation
On this page
The articles use elementary algebra, finite-dimensional linear algebra, derivatives and discrete probability. This appendix defines the shared operations and conventions. Article-specific symbols are defined at their point of use. The following recurring symbols are reserved across the series.
| Symbol | Meaning |
|---|---|
| Network depth, measured in successive layers | |
| , | Loss objective and single-example or single-position loss |
| Optimization objective, including regularization when specified | |
| Scalar image height and width | |
| , | Weight matrix, including a layer index where needed |
| Attention projection matrices | |
| Weight vector or scalar, according to the stated shape | |
| , | Vocabulary size and attention value matrix |
Scalar image width and matrix weights differ typographically as well as by shape. Other matrices retain their conventional letters, such as . Shapes are stated locally whenever an object is introduced.
Displayed equations carry no terminal sentence punctuation. The preceding prose ends with a colon, and the following prose starts a new sentence. Internal commas separating mathematical declarations remain part of the notation.
1. Vector and matrix shapes
A single vector is a column , and denotes its transpose. For , the product is defined coordinatewise by:
For and , matrix multiplication follows the same summation rule:
Inner dimensions must match. Matrix multiplication is generally not commutative. The map is affine and becomes linear when . The notation denotes the identity matrix, an all-ones vector, and coordinatewise multiplication.
Batch feature vectors are arranged as rows of , giving predictions . Image tensors and sequence tensors require separately declared axis conventions.
2. Norm, rank and null space
The Euclidean norm and its square are:
Matrix rank is the maximum number of linearly independent columns. The null space is . For a matrix with columns, rank–nullity states .
A brief justification starts with a basis of the null space and extends it to a basis of the input space. The images of the added basis vectors span the image of and are independent: a dependence among those images would place a nontrivial combination of the added vectors in the null space, contradicting independence of the extended basis. Their count is therefore the rank, giving the stated identity.
Orthogonality means zero dot product. For orthogonal , expansion of eliminates the cross terms and gives .
3. Derivatives, gradients and the chain rule
A scalar derivative is defined by the difference-quotient limit:
For the square function, , which tends to . A partial derivative varies one coordinate while holding the others fixed. The column of these derivatives is the gradient .
For , the Jacobian entry in row , column is . A scalar loss obeys:
Backpropagation applies this rule in reverse computational order. A derivative describes local sensitivity; a finite parameter update additionally specifies a step size.
4. Exponentials and logarithms
The natural exponential is , with inverse for . The identity converts probability products to sums. Derivatives and , combined with the chain rule, give the softmax loss derivative.
Specifically, for , differentiation of the second term gives , while the first contributes . Their sum is .
5. Probability, expectation and variance
A finite distribution satisfies and . For values , its expectation is . This probability-weighted quantity differs from an empirical average of observed samples.
For , conditional probability is . Rearranging gives ; repeated application factorizes a sequence probability without requiring independence between positions.
Variance is . Independence is an additional assumption used, for example, to remove covariance terms in the variance of a sample average.
6. Indicators, indexing and approximations
The indicator is one when its condition holds and zero otherwise. The floor is the greatest integer no larger than . Array positions and token IDs use zero-based indexing where stated; sample and layer indices may start at one.
The notation means as a nonzero displacement approaches zero. It specifies a local remainder and does not justify ignoring that term for a large step. Rounded decimals are approximations; exact fractions and expressions are retained where relevant to a derivation.