From fixed features to trainable convolutional networks
Make feature extraction part of the optimization problem.
On this page
A fixed edge filter encodes a particular choice about which local pattern matters. For a prediction task, the useful patterns may depend on the data. Making the filter coefficients trainable allows the final task error to influence the representation itself, connecting feature extraction to the same optimization process as prediction.
The transition from a fixed-feature pipeline to an end-to-end convolutional network changes which parts of the computation are optimized. The distinction can be stated directly through the training objective: the predictor may be fitted on a fixed representation, or the representation and predictor may be fitted jointly.
1. Two optimization problems
For a fixed representation , the prediction parameters satisfy the optimization problem:
For a representation with trainable parameters, the objective becomes:
The second formulation extends the adjustable parameter set from to . Both pipelines learn prediction parameters; the second also learns the representation. A pretrained deep representation can instead be held fixed, producing a hybrid pipeline. [1]
2. Multichannel convolution
Let the input be , with kernel and bias:
For stride one and explicitly padded input , the preactivation at spatial position and output channel is:
Each output channel combines all input channels in this ordinary, ungrouped convolution. A filter thus spans both the spatial neighborhood and the input-channel axis. The activation is subsequently . [2]
With one bias per output channel, the parameter count is . A layer mapping three channels to six has parameters, independent of the number of spatial positions at which those parameters are reused.
3. Gradients of shared coefficients
Define the upstream preactivation derivative as . Applying the chain rule to a coefficient reused at multiple positions gives:
The spatial sum arises from weight sharing. A batch introduces another sum over samples. If the loss already averages over samples, its scaling is included in and should not be applied a second time.
For ReLU, , with zero conventionally assigned as the backward multiplier at the nondifferentiable zero point. A trainable kernel can then be updated by or another specified optimizer. The target is the task loss, not a prescribed kernel such as Sobel. [3]
4. A single-coefficient update
The scalar case isolates the update mechanism. Let , target , and loss . At , the output is two, the residual is , and:
With learning rate , the coefficient becomes , the output becomes , and the loss falls from eight to . This calculation uses one coefficient at one position; the shared-kernel formula above adds contributions from every position affected by a coefficient.
5. Feature learning and architectural constraints
Repeated trainable convolutions and nonlinearities define a learned representation . The training objective determines which responses are useful, so learned channels can combine patterns more complex than the fixed edge response of the preceding article.
A pretrained may also be fixed while only is updated. This remains a CNN-based feature pipeline. Freezing means that the current optimizer does not modify a parameter; it does not mean that changing that parameter would have no mathematical effect on the output or loss.
End-to-end learning fits feature extraction and prediction within a chosen model structure, objective and training configuration. [1]