Images as tensors: representation and numerical transformations
Read spatial axes, channels and numerical encodings.
On this page
Before a network can process an image, spatial positions and channel values must be assigned explicit numerical coordinates. The distinction matters immediately: changing units, averaging channels and rearranging entries can produce similar-looking arrays while preserving different properties. A six-pixel example makes those effects directly calculable.
A digital image can be represented as a finite array of numerical channel values. This representation separates spatial position, channel identity and numerical encoding. Operations such as normalization, quantization, channel reduction and flattening affect different properties and should not be treated as interchangeable forms of preprocessing.
1. Spatial positions and channels
For height , width and channel count , the mathematical representation is:
The zero-based indices satisfy , and . This notation places channels last. A batch can be arranged as or ; the axis convention must be stated explicitly. Integer storage is a restricted numerical encoding of the real-valued representation.
An RGB example with two rows and three columns is:
| Position | R | G | B |
|---|---|---|---|
| 255 | 0 | 0 | |
| 0 | 255 | 0 | |
| 0 | 0 | 255 | |
| 0 | 0 | 0 | |
| 255 | 255 | 255 | |
| 128 | 128 | 128 |
There are six spatial positions and 18 scalar entries. A pixel is a position with associated channel values, not necessarily a single scalar. The number and meaning of channels depend on the image mode. [1]
2. Scaling and normalization
For an 8-bit channel value , division by 255 gives . For example, . This changes the numerical scale while retaining the same spatial and channel axes.
Channelwise standardization can then use a mean and a positive scale :
When these statistics are estimated from data, they belong to the preprocessing fitted on the training portion. [2] For known , the transformation is invertible in exact arithmetic through . Finite precision or later clipping can introduce additional effects not represented by that algebraic identity.
3. Quantization error
For a normalized value , uniform quantization to levels can be defined by:
Nearest-integer rounding introduces at most half a quantization step, so . For , the bound is . The argument assumes that is in the stated interval. Clipping an out-of-range value can introduce a larger error.
Quantization is therefore different from an exact change of scale: multiple nearby input values may receive the same stored code, making the operation non-injective.
4. Channel reduction
A linear reduction from three channels to one has the form:
For the arithmetic example , each of , and maps to 85. These coefficients define a simple average, not a standard luminance conversion.
The direction lies in this map's null space, since increasing the first channel and decreasing the second by the same amount preserves the average. Whenever both inputs remain in the permitted range, distinct channel triples can therefore produce the same output. The original triple cannot be uniquely recovered from that output alone.
5. Flattening and axis order
Flattening changes the indexing scheme without changing the values. Under row-major, channels-last ordering, the vector index is:
The inverse mapping is:
These expressions establish a one-to-one correspondence between the original coordinates and vector positions. Thus flattening is reversible when the original shape and order are known. It does not have the same information-loss mechanism as averaging channels or quantizing values. A subsequent projection to a shorter vector is a separate operation requiring its own analysis.