AI coursesMain page
Back to the library

Images as tensors: representation and numerical transformations

Read spatial axes, channels and numerical encodings.

On this page
  1. Spatial positions and channels
  2. Scaling and normalization
  3. Quantization error
  4. Channel reduction
  5. Flattening and axis order
  6. References

Before a network can process an image, spatial positions and channel values must be assigned explicit numerical coordinates. The distinction matters immediately: changing units, averaging channels and rearranging entries can produce similar-looking arrays while preserving different properties. A six-pixel example makes those effects directly calculable.

A digital image can be represented as a finite array of numerical channel values. This representation separates spatial position, channel identity and numerical encoding. Operations such as normalization, quantization, channel reduction and flattening affect different properties and should not be treated as interchangeable forms of preprocessing.

1. Spatial positions and channels

For height HH, width WW and channel count CC, the mathematical representation is:

X∈RH×W×C,Xi,j,cX\in\mathbb R^{H\times W\times C},\qquad X_{i,j,c}

The zero-based indices satisfy 0≤i<H0\le i<H, 0≤j<W0\le j<W and 0≤c<C0\le c<C. This notation places channels last. A batch can be arranged as (N,H,W,C)(N,H,W,C) or (N,C,H,W)(N,C,H,W); the axis convention must be stated explicitly. Integer storage is a restricted numerical encoding of the real-valued representation.

An RGB example with two rows and three columns is:

Position (i,j)(i,j) R G B
(0,0)(0,0) 255 0 0
(0,1)(0,1) 0 255 0
(0,2)(0,2) 0 0 255
(1,0)(1,0) 0 0 0
(1,1)(1,1) 255 255 255
(1,2)(1,2) 128 128 128

There are six spatial positions and 18 scalar entries. A pixel is a position with associated channel values, not necessarily a single scalar. The number and meaning of channels depend on the image mode. [1]

2. Scaling and normalization

For an 8-bit channel value Xi,j,c∈{0,…,255}X_{i,j,c}\in\{0,\ldots,255\}, division by 255 gives X~i,j,c∈[0,1]\widetilde X_{i,j,c}\in[0,1]. For example, 128/255≈0.501961128/255\approx0.501961. This changes the numerical scale while retaining the same spatial and channel axes.

Channelwise standardization can then use a mean μc\mu_c and a positive scale scs_c:

Zi,j,c=X~i,j,c−μcscZ_{i,j,c}=\frac{\widetilde X_{i,j,c}-\mu_c}{s_c}

When these statistics are estimated from data, they belong to the preprocessing fitted on the training portion. [2] For known μc,sc\mu_c,s_c, the transformation is invertible in exact arithmetic through X~i,j,c=scZi,j,c+μc\widetilde X_{i,j,c}=s_cZ_{i,j,c}+\mu_c. Finite precision or later clipping can introduce additional effects not represented by that algebraic identity.

3. Quantization error

For a normalized value u∈[0,1]u\in[0,1], uniform quantization to M≥2M\ge2 levels can be defined by:

q=round⁡((M−1)u),u^=qM−1q=\operatorname{round}((M-1)u),\qquad \hat u=\frac{q}{M-1}

Nearest-integer rounding introduces at most half a quantization step, so ∣u^−u∣≤1/[2(M−1)]|\hat u-u|\le1/[2(M-1)]. For M=256M=256, the bound is 1/5101/510. The argument assumes that uu is in the stated interval. Clipping an out-of-range value can introduce a larger error.

Quantization is therefore different from an exact change of scale: multiple nearby input values may receive the same stored code, making the operation non-injective.

4. Channel reduction

A linear reduction from three channels to one has the form:

Yi,j=aRXi,j,0+aGXi,j,1+aBXi,j,2Y_{i,j}=a_R X_{i,j,0}+a_G X_{i,j,1}+a_B X_{i,j,2}

For the arithmetic example aR=aG=aB=1/3a_R=a_G=a_B=1/3, each of (255,0,0)(255,0,0), (0,255,0)(0,255,0) and (0,0,255)(0,0,255) maps to 85. These coefficients define a simple average, not a standard luminance conversion.

The direction (1,−1,0)⊤(1,-1,0)^\top lies in this map's null space, since increasing the first channel and decreasing the second by the same amount preserves the average. Whenever both inputs remain in the permitted range, distinct channel triples can therefore produce the same output. The original triple cannot be uniquely recovered from that output alone.

5. Flattening and axis order

Flattening changes the indexing scheme without changing the values. Under row-major, channels-last ordering, the vector index is:

k=(iW+j)C+c,0≤k<HWCk=(iW+j)C+c,\qquad 0\le k<HWC

The inverse mapping is:

c=k mod C,j=⌊k/C⌋ mod W,i=⌊k/(WC)⌋c=k\bmod C,\quad j=\lfloor k/C\rfloor\bmod W,\quad i=\lfloor k/(WC)\rfloor

These expressions establish a one-to-one correspondence between the original coordinates and vector positions. Thus flattening is reversible when the original shape and order are known. It does not have the same information-loss mechanism as averaging channels or quantizing values. A subsequent projection to a shorter vector is a separate operation requiring its own analysis.

References