Compression, reconstruction and channel bottlenecks
Distinguish a smaller code from a cheaper operation.
On this page
A narrow intermediate representation can serve two different goals: force a reconstruction model to use a limited code, or reduce the cost of an expensive operation. The later expansion restores an output shape required by the task or surrounding layers. Whether it can also restore the original values depends on what the earlier reduction retained.
Compression followed by expansion serves different purposes in different architectures. A smaller spatial grid, fewer channels and a lower-dimensional vector impose different constraints. The relevant question is which representation is reduced, what the task requires it to preserve, and what information remains available to the subsequent expansion.
1. Encoding and reconstruction
An encoder maps to , and a decoder produces . When , the code forms a vector bottleneck. For samples, a squared reconstruction objective is:
The objective rewards codes that preserve what the decoder needs for reconstruction. Which distinctions survive depends on the code capacity, model family and data distribution. [1]
2. Non-invertibility of a linear bottleneck
For example, maps both and to one. Decoding with returns for either input, giving squared errors two and zero respectively.
For a linear encoder with and , rank–nullity gives:
A nonzero vector therefore exists with , implying . Different inputs in can have the same code. No decoder receiving only that code can distinguish all such inputs. For a linear decoder , the additional inequality rules out .
The proof concerns a linear encoder on the entire input space. A restricted data set may lie in a subspace small enough for exact reconstruction; orthogonal projection provides an explicit example. [1]
3. Orthogonal projection and reconstruction error
Let have orthonormal columns, so . Encoding with and decoding with gives:
Since , expansion of the squared norm yields:
The residual is the component outside the selected subspace. Inputs in the subspace are reconstructed exactly, while the squared norm of the discarded component gives the reconstruction error.
4. Spatial downsampling and upsampling
Reducing spatial resolution can lower the cost of subsequent spatial operations and enlarge the spacing between their input dependencies. It can also remove fine positional detail. These effects concern the spatial axes and are not equivalent to reducing channel width at every position.
Upsampling increases spatial resolution through replication, interpolation or a learned mapping. It is an inverse only if composition with the preceding operation recovers every input in the stated domain. The averaging example shows that increasing output size does not resolve an ambiguity already introduced by compression.
Encoder-to-decoder skip connections provide representations that bypass a bottleneck. Concatenation adds channels, whereas addition combines corresponding entries of matching shapes. These paths make earlier spatial detail available to the decoder. [2]
5. Channel bottlenecks and computational cost
Compare a convolution from 64 to 64 channels with three layers: , , and . Keeping spatial sizes fixed and excluding biases gives:
Under these assumptions, the same counts describe MACs per output position. The bottleneck reduces the width of the expensive spatial operation before restoring the external channel size. The complete block is assessed together with its nonlinearities and residual branches. [3]
Reconstruction constraints, spatial resolution and channel-operation cost therefore provide distinct explanations for compression followed by expansion. None establishes a universal requirement that networks must use this pattern.