AI coursesMain page
Back to the library

From fixed features to trainable convolutional networks

Make feature extraction part of the optimization problem.

On this page
  1. Two optimization problems
  2. Multichannel convolution
  3. Gradients of shared coefficients
  4. A single-coefficient update
  5. Feature learning and architectural constraints
  6. References

A fixed edge filter encodes a particular choice about which local pattern matters. For a prediction task, the useful patterns may depend on the data. Making the filter coefficients trainable allows the final task error to influence the representation itself, connecting feature extraction to the same optimization process as prediction.

The transition from a fixed-feature pipeline to an end-to-end convolutional network changes which parts of the computation are optimized. The distinction can be stated directly through the training objective: the predictor may be fitted on a fixed representation, or the representation and predictor may be fitted jointly.

1. Two optimization problems

For a fixed representation ϕ\phi, the prediction parameters satisfy the optimization problem:

β∗∈arg⁡min⁡β1n∑iℓ(gβ(ϕ(xi)),yi)\beta^*\in\arg\min_\beta \frac1n\sum_i\ell(g_\beta(\phi(x_i)),y_i)

For a representation ϕα\phi_\alpha with trainable parameters, the objective becomes:

(α∗,β∗)∈arg⁡min⁡α,β1n∑iℓ(gβ(ϕα(xi)),yi)(\alpha^*,\beta^*)\in\arg\min_{\alpha,\beta} \frac1n\sum_i\ell(g_\beta(\phi_\alpha(x_i)),y_i)

The second formulation extends the adjustable parameter set from β\beta to (α,β)(\alpha,\beta). Both pipelines learn prediction parameters; the second also learns the representation. A pretrained deep representation can instead be held fixed, producing a hybrid pipeline. [1]

2. Multichannel convolution

Let the input be X∈RH×W×CinX\in\mathbb R^{H\times W\times C_{\rm in}}, with kernel and bias:

K∈Rkh×kw×Cin×Cout,b∈RCoutK\in\mathbb R^{k_h\times k_w\times C_{\rm in}\times C_{\rm out}}, \qquad b\in\mathbb R^{C_{\rm out}}

For stride one and explicitly padded input X~\widetilde X, the preactivation at spatial position (i,j)(i,j) and output channel oo is:

Zi,j,o=bo+∑u=0kh−1∑v=0kw−1∑c=0Cin−1Ku,v,c,oX~i+u,j+v,cZ_{i,j,o}=b_o+ \sum_{u=0}^{k_h-1}\sum_{v=0}^{k_w-1} \sum_{c=0}^{C_{\rm in}-1} K_{u,v,c,o}\widetilde X_{i+u,j+v,c}

Each output channel combines all input channels in this ordinary, ungrouped convolution. A filter thus spans both the spatial neighborhood and the input-channel axis. The activation is subsequently Ai,j,o=σ(Zi,j,o)A_{i,j,o}=\sigma(Z_{i,j,o}). [2]

With one bias per output channel, the parameter count is khkwCinCout+Coutk_hk_wC_{\rm in}C_{\rm out}+C_{\rm out}. A 3×33\times3 layer mapping three channels to six has 162+6=168162+6=168 parameters, independent of the number of spatial positions at which those parameters are reused.

3. Gradients of shared coefficients

Define the upstream preactivation derivative as δi,j,o=∂L/∂Zi,j,o\delta_{i,j,o}=\partial\mathcal L/\partial Z_{i,j,o}. Applying the chain rule to a coefficient reused at multiple positions gives:

∂L∂Ku,v,c,o=∑i,jδi,j,oX~i+u,j+v,c,∂L∂bo=∑i,jδi,j,o\frac{\partial\mathcal L}{\partial K_{u,v,c,o}} =\sum_{i,j}\delta_{i,j,o}\widetilde X_{i+u,j+v,c}, \qquad \frac{\partial\mathcal L}{\partial b_o}=\sum_{i,j}\delta_{i,j,o}

The spatial sum arises from weight sharing. A batch introduces another sum over samples. If the loss already averages over samples, its scaling is included in δ\delta and should not be applied a second time.

For ReLU, δ=(∂L/∂A)⊙1[Z>0]\delta=(\partial\mathcal L/\partial A)\odot\mathbf1[Z>0], with zero conventionally assigned as the backward multiplier at the nondifferentiable zero point. A trainable kernel can then be updated by K←K−η∇KLK\leftarrow K-\eta\nabla_K\mathcal L or another specified optimizer. The target is the task loss, not a prescribed kernel such as Sobel. [3]

4. A single-coefficient update

The scalar case z=kxz=kx isolates the update mechanism. Let x=2x=2, target y=6y=6, and loss ℓ=(z−y)2/2\ell=(z-y)^2/2. At k=1k=1, the output is two, the residual is −4-4, and:

∂ℓ∂k=(kx−y)x=−8\frac{\partial\ell}{\partial k}=(kx-y)x=-8

With learning rate η=0.1\eta=0.1, the coefficient becomes k′=1.8k'=1.8, the output becomes z′=3.6z'=3.6, and the loss falls from eight to 2.882.88. This calculation uses one coefficient at one position; the shared-kernel formula above adds contributions from every position affected by a coefficient.

5. Feature learning and architectural constraints

Repeated trainable convolutions and nonlinearities define a learned representation ϕα\phi_\alpha. The training objective determines which responses are useful, so learned channels can combine patterns more complex than the fixed edge response of the preceding article.

A pretrained α\alpha may also be fixed while only β\beta is updated. This remains a CNN-based feature pipeline. Freezing means that the current optimizer does not modify a parameter; it does not mean that changing that parameter would have no mathematical effect on the output or loss.

End-to-end learning fits feature extraction and prediction within a chosen model structure, objective and training configuration. [1]

References