AI coursesMain page
Back to the library

Expansion, nonlinear transformation and projection

Build a nonlinear function through a wider intermediate layer.

On this page
  1. Expansion and projection in an MLP
  2. An exact two–four–two construction
  3. Expressivity and local rank
  4. Connection to Transformer feed-forward layers
  5. Computational tradeoff
  6. References

A wider hidden layer can compute several responses to the same input and then combine them into a compact output. For example, separate responses to positive and negative values can be recombined to produce an absolute value. This simple construction explains the role of expansion before any reference to a Transformer is needed.

An expanded hidden layer provides additional coordinates for an intermediate nonlinear transformation while preserving a specified input and output width. Its effect is determined by the activation and recombination of those coordinates, rather than by an increase in coordinate count alone.

1. Expansion and projection in an MLP

For a column input x∈Rdx\in\mathbb R^d and hidden width m>dm>d, a two-layer multilayer perceptron (MLP) can be written as:

h=ReLU⁡(Ax+a),y=Bh+bh=\operatorname{ReLU}(Ax+a),\qquad y=Bh+b

The parameter shapes are:

A∈Rm×d,a∈Rm,B∈Rd×m,b∈RdA\in\mathbb R^{m\times d},\quad a\in\mathbb R^m,\quad B\in\mathbb R^{d\times m},\quad b\in\mathbb R^d

ReLU acts coordinatewise as ReLU⁡(t)=max⁡(0,t)\operatorname{ReLU}(t)=\max(0,t). The hidden representation has mm coordinates and the output returns to dd. Here “projection” means a learned linear map; it need not be an orthogonal projection. Including both biases gives 2dm+m+d2dm+m+d parameters. [1]

Without the activation, the function reduces to y=BAx+Ba+by=BAx+Ba+b, a single affine map. Consequently, expansion alone does not establish a larger family of nonlinear functions. The activation is essential to the construction considered here.

2. An exact two–four–two construction

Take zero biases and matrices:

A=(10−10010−1),B=(11000011)A=\begin{pmatrix}1&0\\-1&0\\0&1\\0&-1\end{pmatrix},\qquad B=\begin{pmatrix}1&1&0&0\\0&0&1&1\end{pmatrix}

The hidden and output vectors are then:

h=(max⁡(0,x1)max⁡(0,−x1)max⁡(0,x2)max⁡(0,−x2)),y=(∣x1∣∣x2∣)h=\begin{pmatrix}\max(0,x_1)\\\max(0,-x_1)\\\max(0,x_2)\\\max(0,-x_2)\end{pmatrix},\qquad y=\begin{pmatrix}|x_1|\\|x_2|\end{pmatrix}

For each coordinate, one hidden unit retains its positive part and another retains the positive part of its negation. Their sum is the absolute value. For x=(−2,3)⊤x=(-2,3)^\top:

Ax=(−2,2,3,−3)⊤,h=(0,2,3,0)⊤,y=(2,3)⊤Ax=(-2,2,3,-3)^\top,\quad h=(0,2,3,0)^\top,\quad y=(2,3)^\top

The architecture has 16 matrix entries and six bias slots. The explicit construction above sets all biases to zero and specifies the matrix values directly.

Removing ReLU gives BA=0BA=0, so every input would produce zero. The nonlinear activation thus changes the represented function in a directly verifiable way.

3. Expressivity and local rank

No single affine function at+cat+c represents ∣t∣|t| on the real line. The values at zero and one require c=0c=0 and a=1a=1, while the value at minus one requires a=−1a=-1, a contradiction.

Within a region with fixed ReLU activation signs, the MLP is affine. Away from zero preactivations, its Jacobian is:

Jy(x)=B diag⁡ ⁣(1[Ax+a>0])AJ_y(x)=B\,\operatorname{diag}\!\left(\mathbf1[Ax+a>0]\right)A

The diagonal entries select active units. Crossing an activation boundary can change the local affine map. A larger mm provides more intermediate units from which training can form useful activation patterns.

The Jacobian still has rank at most dd. Additional coordinates computed from the same input do not create independent observations. In the explicit construction, x1=h1−h2x_1=h_1-h_2 and x2=h3−h4x_2=h_3-h_4, so the hidden representation retains the input. The final sums lose the signs. Information retention therefore depends on the particular mapping, not on the words “expansion” or “projection” alone. [1]

4. Connection to Transformer feed-forward layers

The original Transformer uses an expand–activate–project structure in its position-wise feed-forward network. The same function is applied separately at each sequence position with shared parameters. This differs from attention, which combines representations across positions. [2]

For the original configuration d=512d=512, m=2048m=2048, the parameter count including biases is:

2⋅512⋅2048+2048+512=20997122\cdot512\cdot2048+2048+512=2099712

This count includes the feed-forward layer alone. The original fourfold width is one design choice; the local mechanism remains expansion, nonlinear activation and projection back to the required output width.

5. Computational tradeoff

Increasing mm increases hidden-activation storage and the arithmetic cost of the matrix products. In return, the block can form and recombine more nonlinear intermediate responses while retaining an output shape compatible with surrounding layers.

Whether that tradeoff is beneficial depends on the task, optimization procedure and held-out evaluation. Neither width nor parameter count alone determines accuracy. Expansion followed by projection is therefore best understood as a structured nonlinear computation with an explicit cost, rather than as a universal rule for adding or removing information. [1]

References