AI coursesMain page
Back to the library

Mathematical notation

On this page
  1. Vector and matrix shapes
  2. Norm, rank and null space
  3. Derivatives, gradients and the chain rule
  4. Exponentials and logarithms
  5. Probability, expectation and variance
  6. Indicators, indexing and approximations

The articles use elementary algebra, finite-dimensional linear algebra, derivatives and discrete probability. This appendix defines the shared operations and conventions. Article-specific symbols are defined at their point of use. The following recurring symbols are reserved across the series.

Symbol Meaning
LnetL_{\mathrm{net}} Network depth, measured in successive layers
L\mathcal L, ℓ\ell Loss objective and single-example or single-position loss
JJ Optimization objective, including regularization when specified
H,WH,W Scalar image height and width
W\mathbf{W}, Wl\mathbf{W}_l Weight matrix, including a layer index where needed
WQ,WK,WV\mathbf{W}_Q,\mathbf{W}_K,\mathbf{W}_V Attention projection matrices
ww Weight vector or scalar, according to the stated shape
VV, VhV_h Vocabulary size and attention value matrix

Scalar image width WW and matrix weights W\mathbf{W} differ typographically as well as by shape. Other matrices retain their conventional letters, such as A,B,EA,B,E. Shapes are stated locally whenever an object is introduced.

Displayed equations carry no terminal sentence punctuation. The preceding prose ends with a colon, and the following prose starts a new sentence. Internal commas separating mathematical declarations remain part of the notation.

1. Vector and matrix shapes

A single vector is a column x∈Rdx\in\mathbb R^d, and x⊤x^\top denotes its transpose. For A∈Rm×dA\in\mathbb R^{m\times d}, the product y=Ax∈Rmy=Ax\in\mathbb R^m is defined coordinatewise by:

yi=∑j=1dAijxjy_i=\sum_{j=1}^d A_{ij}x_j

For A∈Rm×dA\in\mathbb R^{m\times d} and B∈Rd×nB\in\mathbb R^{d\times n}, matrix multiplication follows the same summation rule:

(AB)ik=∑j=1dAijBjk,AB∈Rm×n(AB)_{ik}=\sum_{j=1}^d A_{ij}B_{jk},\qquad AB\in\mathbb R^{m\times n}

Inner dimensions must match. Matrix multiplication is generally not commutative. The map Ax+bAx+b is affine and becomes linear when b=0b=0. The notation IdI_d denotes the d×dd\times d identity matrix, 1n\mathbf1_n an all-ones vector, and u⊙vu\odot v coordinatewise multiplication.

Batch feature vectors are arranged as rows of X∈Rn×dX\in\mathbb R^{n\times d}, giving predictions Xw+b1nXw+b\mathbf1_n. Image tensors and sequence tensors require separately declared axis conventions.

2. Norm, rank and null space

The Euclidean norm and its square are:

∥x∥2=∑jxj2,∥x∥22=x⊤x\|x\|_2=\sqrt{\sum_jx_j^2},\qquad \|x\|_2^2=x^\top x

Matrix rank is the maximum number of linearly independent columns. The null space is ker⁡A={v:Av=0}\ker A=\{v:Av=0\}. For a matrix with dd columns, rank–nullity states d=rank⁡(A)+dim⁡ker⁡Ad=\operatorname{rank}(A)+\dim\ker A.

A brief justification starts with a basis of the null space and extends it to a basis of the input space. The images of the added basis vectors span the image of AA and are independent: a dependence among those images would place a nontrivial combination of the added vectors in the null space, contradicting independence of the extended basis. Their count is therefore the rank, giving the stated identity.

Orthogonality means zero dot product. For orthogonal u,vu,v, expansion of (u+v)⊤(u+v)(u+v)^\top(u+v) eliminates the cross terms and gives ∥u+v∥22=∥u∥22+∥v∥22\|u+v\|_2^2=\|u\|_2^2+\|v\|_2^2.

3. Derivatives, gradients and the chain rule

A scalar derivative is defined by the difference-quotient limit:

f′(t)=lim⁡h→0f(t+h)−f(t)hf'(t)=\lim_{h\to0}\frac{f(t+h)-f(t)}h

For the square function, ((t+h)2−t2)/h=2t+h((t+h)^2-t^2)/h=2t+h, which tends to 2t2t. A partial derivative ∂J/∂θj\partial J/\partial\theta_j varies one coordinate while holding the others fixed. The column of these derivatives is the gradient ∇J\nabla J.

For y=f(x)y=f(x), the Jacobian entry in row ii, column jj is ∂yi/∂xj\partial y_i/\partial x_j. A scalar loss obeys:

∂L∂xj=∑i∂L∂yi∂yi∂xj,∇xL=Jf(x)⊤∇yL\frac{\partial \mathcal L}{\partial x_j}=\sum_i\frac{\partial \mathcal L}{\partial y_i}\frac{\partial y_i}{\partial x_j},\qquad \nabla_xL=J_f(x)^\top\nabla_yL

Backpropagation applies this rule in reverse computational order. A derivative describes local sensitivity; a finite parameter update additionally specifies a step size.

4. Exponentials and logarithms

The natural exponential is ete^t, with inverse log⁡t\log t for t>0t>0. The identity log⁡(ab)=log⁡a+log⁡b\log(ab)=\log a+\log b converts probability products to sums. Derivatives (et)′=et(e^t)'=e^t and (log⁡t)′=1/t(\log t)'=1/t, combined with the chain rule, give the softmax loss derivative.

Specifically, for ℓ=−zy+log⁡∑jezj\ell=-z_y+\log\sum_j e^{z_j}, differentiation of the second term gives ezk/∑jezje^{z_k}/\sum_j e^{z_j}, while the first contributes −1[k=y]-\mathbf1[k=y]. Their sum is pk−1[k=y]p_k-\mathbf1[k=y].

5. Probability, expectation and variance

A finite distribution satisfies pi≥0p_i\ge0 and ∑ipi=1\sum_i p_i=1. For values aia_i, its expectation is E[a]=∑ipiai\mathbb E[a]=\sum_i p_i a_i. This probability-weighted quantity differs from an empirical average of observed samples.

For p(b)>0p(b)>0, conditional probability is p(a∣b)=p(a,b)/p(b)p(a\mid b)=p(a,b)/p(b). Rearranging gives p(a,b)=p(a∣b)p(b)p(a,b)=p(a\mid b)p(b); repeated application factorizes a sequence probability without requiring independence between positions.

Variance is Var⁡(Z)=E[(Z−EZ)2]=E[Z2]−(EZ)2\operatorname{Var}(Z)=\mathbb E[(Z-\mathbb EZ)^2]=\mathbb E[Z^2]-(\mathbb EZ)^2. Independence is an additional assumption used, for example, to remove covariance terms in the variance of a sample average.

6. Indicators, indexing and approximations

The indicator 1[condition]\mathbf1[\text{condition}] is one when its condition holds and zero otherwise. The floor ⌊t⌋\lfloor t\rfloor is the greatest integer no larger than tt. Array positions and token IDs use zero-based indexing where stated; sample and layer indices may start at one.

The notation r(Δ)=o(∥Δ∥)r(\Delta)=o(\|\Delta\|) means r(Δ)/∥Δ∥→0r(\Delta)/\|\Delta\|\to0 as a nonzero displacement approaches zero. It specifies a local remainder and does not justify ignoring that term for a large step. Rounded decimals are approximations; exact fractions and expressions are retained where relevant to a derivation.