Tensor shapes, parameter counts and computational cost
Calculate output shapes, parameter counts and computation.
On this page
A shape calculation answers practical design questions before a layer is run: whether two blocks can connect, how much activation storage is needed, and how much arithmetic a convolution requires. Naming the axes first turns a change such as eight-by-eight-by-three to four-by-four-by-six into several measurable decisions rather than one vague change in dimension.
A tensor shape records the sizes of named axes. It does not by itself specify the number of trainable parameters, the computational cost or the information preserved by a transformation. These quantities require separate definitions and calculations.
1. Axes, vector dimension and matrix rank
For a channels-last image, denotes height, width and channel count. A batch has shape , whereas the channels-first convention is . Axis meanings cannot be recovered reliably from an unlabeled list of numbers.
The number of tensor axes is sometimes called tensor rank in array libraries. This differs from matrix rank, which counts linearly independent columns. Vector dimension instead counts coordinates in a vector. The term “dimension” must therefore be interpreted in its stated context.
An image of shape contains scalar entries. Flattening gives a vector in without discarding entries if their order is preserved. Entry count alone, however, is not a measure of information content.
2. Derivation of convolution output size
Consider an input axis of length , symmetric padding , kernel length , dilation and stride . Assume positive integer and nonnegative integer . The span from the first to the last kernel sample gives the effective kernel size:
In the padded input, window starts are . The last window fits when . The number of valid starts is consequently:
The same argument applies to width. At least one window must fit; asymmetric padding replaces with the sum of the two padding amounts. The floor cannot be omitted when the available span is not divisible by the stride. [1]
3. Activations, parameters and operations
For , , , , , and :
The transformation is : height and width halve, while channel count doubles. Each sample has 192 input entries and 96 output entries. Including one bias per output channel, the shared parameter count is:
Counting one multiply–accumulate (MAC) for each kernel coefficient at each output position, and excluding bias additions, gives:
A batch of size multiplies the activation count and MAC count by , but leaves the parameter count at 168. The operation-counting convention matters: one MAC may be reported as two floating-point operations.
4. Grouped and depthwise convolution
For groups, both input and output channel counts must be divisible by . Each output channel accesses input channels, yielding:
Depthwise convolution corresponds to , with an integer multiple of . Grouping restricts channel connectivity; the lower count does not preserve every interaction available in an ordinary dense convolution. [1]
5. Receptive-field span
Let denote the theoretical receptive-field span at layer , and the spacing between neighboring outputs measured in original input coordinates. Starting from :
The additional span contains gaps, each of size . For two kernel-size-three layers with dilation one and strides two then one, the results are and .
The span describes a possible input dependency, not equal influence from every position. Padding also requires distinguishing original data from extended coordinates at boundaries. A complete shape analysis therefore states the changing axes, activation and parameter counts, operation count, and possible spatial dependencies separately. [2]