Expansion, nonlinear transformation and projection
Build a nonlinear function through a wider intermediate layer.
On this page
A wider hidden layer can compute several responses to the same input and then combine them into a compact output. For example, separate responses to positive and negative values can be recombined to produce an absolute value. This simple construction explains the role of expansion before any reference to a Transformer is needed.
An expanded hidden layer provides additional coordinates for an intermediate nonlinear transformation while preserving a specified input and output width. Its effect is determined by the activation and recombination of those coordinates, rather than by an increase in coordinate count alone.
1. Expansion and projection in an MLP
For a column input and hidden width , a two-layer multilayer perceptron (MLP) can be written as:
The parameter shapes are:
ReLU acts coordinatewise as . The hidden representation has coordinates and the output returns to . Here “projection” means a learned linear map; it need not be an orthogonal projection. Including both biases gives parameters. [1]
Without the activation, the function reduces to , a single affine map. Consequently, expansion alone does not establish a larger family of nonlinear functions. The activation is essential to the construction considered here.
2. An exact two–four–two construction
Take zero biases and matrices:
The hidden and output vectors are then:
For each coordinate, one hidden unit retains its positive part and another retains the positive part of its negation. Their sum is the absolute value. For :
The architecture has 16 matrix entries and six bias slots. The explicit construction above sets all biases to zero and specifies the matrix values directly.
Removing ReLU gives , so every input would produce zero. The nonlinear activation thus changes the represented function in a directly verifiable way.
3. Expressivity and local rank
No single affine function represents on the real line. The values at zero and one require and , while the value at minus one requires , a contradiction.
Within a region with fixed ReLU activation signs, the MLP is affine. Away from zero preactivations, its Jacobian is:
The diagonal entries select active units. Crossing an activation boundary can change the local affine map. A larger provides more intermediate units from which training can form useful activation patterns.
The Jacobian still has rank at most . Additional coordinates computed from the same input do not create independent observations. In the explicit construction, and , so the hidden representation retains the input. The final sums lose the signs. Information retention therefore depends on the particular mapping, not on the words “expansion” or “projection” alone. [1]
4. Connection to Transformer feed-forward layers
The original Transformer uses an expand–activate–project structure in its position-wise feed-forward network. The same function is applied separately at each sequence position with shared parameters. This differs from attention, which combines representations across positions. [2]
For the original configuration , , the parameter count including biases is:
This count includes the feed-forward layer alone. The original fourfold width is one design choice; the local mechanism remains expansion, nonlinear activation and projection back to the required output width.
5. Computational tradeoff
Increasing increases hidden-activation storage and the arithmetic cost of the matrix products. In return, the block can form and recombine more nonlinear intermediate responses while retaining an output shape compatible with surrounding layers.
Whether that tradeoff is beneficial depends on the task, optimization procedure and held-out evaluation. Neither width nor parameter count alone determines accuracy. Expansion followed by projection is therefore best understood as a structured nonlinear computation with an explicit cost, rather than as a universal rule for adding or removing information. [1]