Features, parameters and hyperparameters
Identify what comes from the input and what training changes.
On this page
When a prediction changes, the cause may be a different input, an updated coefficient, or a different training configuration. Features, parameters and hyperparameters identify these three levels. A small linear model is sufficient to track each level separately before applying the same distinction to a neural network.
Features, parameters and hyperparameters describe different roles in a learning procedure. Features represent an input; parameters determine the model's mapping; hyperparameters configure the training process or model family. Their distinction depends on the specified optimization procedure, rather than on whether a quantity happens to be stored as a number.
1. Features and prediction parameters
For a feature vector , weights and scalar bias , linear regression predicts:
The entries of describe the current sample. The quantities in specify the shared prediction rule. Changing a sample changes without necessarily changing the model; fitting changes using a training objective. [1]
For , and , the prediction is:
Changing the second weight from four to five changes the prediction to . This example isolates a parameter change while leaving the input fixed. A different input under the original weights would instead illustrate a change in activation or output.
2. Batch representation and activations
A dataset containing feature vectors can be arranged as the rows of . Batch prediction is then:
Here is an all-ones vector. Consider the following batch and weights:
With , the result is . There are three model parameters and two output values. Increasing the number of samples increases the number of outputs, but does not increase the number of shared parameters.
The same distinction applies inside a neural network. An activation is an intermediate value produced for a particular input. A weight is a coefficient reused across computations. Both may be arrays, but their mathematical roles differ. [1]
3. Fixed and trainable features
A fixed feature transformation might be . Its third coordinate introduces a nonlinear interaction between the original coordinates. A linear predictor on can therefore represent functions that are not affine in the original .
The transformation is nevertheless fixed if its construction does not change during fitting. By contrast, has adjustable parameters . End-to-end learning can optimize both the representation and the final prediction function. Feature construction and parameter learning are consequently related but distinct decisions.
4. Hyperparameters and model selection
Let collect a configuration such as the learning rate, hidden width, regularization coefficient or training duration. The output of a finite training procedure can be represented as:
This notation includes the procedure and its configuration; it does not assume exact minimization. Given a finite candidate set , validation-based selection has the form:
Hyperparameters are outside the ordinary parameter updates of the specified training run. They need not be selected manually: a separate search or optimization procedure may choose them. The distinction concerns the level of optimization being described. [2]
5. Regularization as an explicit example
For batch targets , a regularized linear-regression objective is:
The first term measures mean squared residual with a factor of . The second penalizes weight magnitude; the bias is not penalized in this definition. The model parameters are , while controls the relative strength of the penalty.
Differentiating the penalty gives , which is added to the unregularized weight gradient. Thus a hyperparameter can directly affect parameter updates without itself being updated by that same rule. A complete specification must identify the adjustable variables, fixed quantities and selection procedure; labels alone are insufficient.