AI coursesMain page
Back to the library

Nonlinearity, depth and residual connections

Explore what nonlinearities and shortcut paths contribute.

On this page
  1. Composition of affine layers
  2. Nonlinearity and piecewise affine functions
  3. Residual blocks
  4. Derivatives and shape compatibility
  5. Depth, width and evaluation
  6. References

Adding layers is useful when their composition can express transformations that a simpler model handles poorly. Nonlinearities create different responses in different input regions; residual paths allow a branch to learn a correction to an existing representation. These two mechanisms explain more about a network than layer count alone.

Network depth, nonlinear activation and residual addition affect different aspects of a model. Their roles can be distinguished by examining the represented function and its derivatives, rather than inferring capability from layer count alone.

1. Composition of affine layers

Consider two shape-compatible affine layers:

h=W1x+b1,y=W2h+b2h=\mathbf{W}_1x+b_1,\qquad y=\mathbf{W}_2h+b_2

Substitution gives:

y=(W2W1)x+(W2b1+b2)y=(\mathbf{W}_2\mathbf{W}_1)x+(\mathbf{W}_2b_1+b_2)

The composition is therefore equivalent, as a function, to a single affine map. The same argument applies recursively to any finite stack containing only affine transformations. Different parameter factorizations can still change optimization behavior, so function equivalence does not imply identical training dynamics. [1]

For scalar layers h=2x+1h=2x+1 and y=3h−2y=3h-2, the combined function is y=6x+1y=6x+1. Increasing the number of such layers alone does not introduce a nonlinear input–output relation.

2. Nonlinearity and piecewise affine functions

Inserting ReLU⁡(z)=max⁡(0,z)\operatorname{ReLU}(z)=\max(0,z) yields:

f(x)=3ReLU⁡(2x+1)−2={−2,x≤−1/2,6x+1,x>−1/2f(x)=3\operatorname{ReLU}(2x+1)-2 =\begin{cases} -2,&x\le-1/2,\\ 6x+1,&x>-1/2 \end{cases}

The two regions have different slopes. A single affine expression cannot represent the function over the whole real line. Direct substitution gives f(−1)=−2f(-1)=-2, f(0)=1f(0)=1 and f(1)=7f(1)=7.

More generally, fixing the active units of a ReLU network replaces each activation with a diagonal mask of zeros and ones. Within the corresponding input region, the network is affine; crossing an activation boundary can select a different affine map. This regional variation provides the nonlinear behavior. [1]

3. Residual blocks

A same-width residual block is defined by:

y=x+Fθ(x),x,y∈Rdy=x+F_\theta(x),\qquad x,y\in\mathbb R^d

The direct path carries xx, while the trainable branch represents an additive correction. For a desired mapping H(x)H(x), that correction can be expressed as H(x)−xH(x)-x. When retaining the input is useful, the branch need not independently reproduce the complete identity mapping. [2]

For x=(2,−1)⊤x=(2,-1)^\top and Fθ(x)=(0.5,0.25)⊤F_\theta(x)=(0.5,0.25)^\top, the output is (2.5,−0.75)⊤(2.5,-0.75)^\top. If the branch returns zero, the block preserves xx, making the identity mapping a particularly simple case.

4. Derivatives and shape compatibility

At differentiable points, let JF=∂F/∂xJ_F=\partial F/\partial x. The input Jacobian and loss gradient are:

Jy=I+JF,∇xL=(I+JF)⊤∇yLJ_y=I+J_F,\qquad \nabla_x\mathcal L=(I+J_F)^\top\nabla_y\mathcal L

The identity term contributes a direct gradient path. Cancellation remains possible: for F(x)=−xF(x)=-x, both the output and its Jacobian are zero. This counterexample identifies the limit of the gradient-path argument.

When input and output widths differ, a projected shortcut can be used:

y=Px+Fθ(x),P∈Rdout×diny=Px+F_\theta(x),\qquad P\in\mathbb R^{d_{\rm out}\times d_{\rm in}}

Both branches must produce the same shape before addition. The projection PP is not an identity map between equal-dimensional spaces, and its effect must be included when analyzing the block. Concatenation is a different operation: it joins coordinates instead of adding corresponding entries. [2]

5. Depth, width and evaluation

Depth counts sequential transformation stages, while width measures the number of components at a stage. Changing either can alter the function family, computational cost and optimization problem. Initialization, normalization and learning rate remain relevant to whether useful parameters can be found.

Three claims must consequently be separated: that an architecture can represent a function, that a training procedure can find suitable parameters, and that the fitted function performs well on the target distribution. The first claim does not prove the other two. A deeper network may provide a useful factorization, but layer count alone is not evidence of better optimization or generalization.

The same distinctions apply to bottlenecks and expanded hidden layers: shape changes must be assessed together with their nonlinear transformations and task-specific evaluation. [2]

References