AI coursesMain page
Back to the library

Features, parameters and hyperparameters

Identify what comes from the input and what training changes.

On this page
  1. Features and prediction parameters
  2. Batch representation and activations
  3. Fixed and trainable features
  4. Hyperparameters and model selection
  5. Regularization as an explicit example
  6. References

When a prediction changes, the cause may be a different input, an updated coefficient, or a different training configuration. Features, parameters and hyperparameters identify these three levels. A small linear model is sufficient to track each level separately before applying the same distinction to a neural network.

Features, parameters and hyperparameters describe different roles in a learning procedure. Features represent an input; parameters determine the model's mapping; hyperparameters configure the training process or model family. Their distinction depends on the specified optimization procedure, rather than on whether a quantity happens to be stored as a number.

1. Features and prediction parameters

For a feature vector x∈Rdx\in\mathbb R^d, weights w∈Rdw\in\mathbb R^d and scalar bias bb, linear regression predicts:

y^=w⊤x+b=∑j=1dwjxj+b\hat y=w^\top x+b=\sum_{j=1}^{d}w_jx_j+b

The entries of xx describe the current sample. The d+1d+1 quantities in (w,b)(w,b) specify the shared prediction rule. Changing a sample changes xx without necessarily changing the model; fitting changes (w,b)(w,b) using a training objective. [1]

For x=(2,−1)⊤x=(2,-1)^\top, w=(3,4)⊤w=(3,4)^\top and b=0.5b=0.5, the prediction is:

y^=3⋅2+4⋅(−1)+0.5=2.5\hat y=3\cdot2+4\cdot(-1)+0.5=2.5

Changing the second weight from four to five changes the prediction to 1.51.5. This example isolates a parameter change while leaving the input fixed. A different input under the original weights would instead illustrate a change in activation or output.

2. Batch representation and activations

A dataset containing nn feature vectors can be arranged as the rows of X∈Rn×dX\in\mathbb R^{n\times d}. Batch prediction is then:

y^=Xw+b1n∈Rn\hat y=Xw+b\mathbf1_n\in\mathbb R^n

Here 1n\mathbf1_n is an all-ones vector. Consider the following batch and weights:

X=[2−101],w=[34]X=\begin{bmatrix}2&-1\\0&1\end{bmatrix},\qquad w=\begin{bmatrix}3\\4\end{bmatrix}

With b=0.5b=0.5, the result is (2.5,4.5)⊤(2.5,4.5)^\top. There are three model parameters and two output values. Increasing the number of samples increases the number of outputs, but does not increase the number of shared parameters.

The same distinction applies inside a neural network. An activation is an intermediate value produced for a particular input. A weight is a coefficient reused across computations. Both may be arrays, but their mathematical roles differ. [1]

3. Fixed and trainable features

A fixed feature transformation might be ϕ(x)=(x1,x2,x1x2)⊤\phi(x)=(x_1,x_2,x_1x_2)^\top. Its third coordinate introduces a nonlinear interaction between the original coordinates. A linear predictor on ϕ(x)\phi(x) can therefore represent functions that are not affine in the original xx.

The transformation is nevertheless fixed if its construction does not change during fitting. By contrast, ϕα(x)\phi_\alpha(x) has adjustable parameters α\alpha. End-to-end learning can optimize both the representation and the final prediction function. Feature construction and parameter learning are consequently related but distinct decisions.

4. Hyperparameters and model selection

Let λ\lambda collect a configuration such as the learning rate, hidden width, regularization coefficient or training duration. The output of a finite training procedure can be represented as:

θλ=Train⁡(Dtrain;λ)\theta_\lambda=\operatorname{Train}(\mathcal D_{\mathrm{train}};\lambda)

This notation includes the procedure and its configuration; it does not assume exact minimization. Given a finite candidate set Λ\Lambda, validation-based selection has the form:

λ∗∈arg⁡min⁡λ∈ΛR^val(θλ)\lambda^*\in\arg\min_{\lambda\in\Lambda} \widehat R_{\mathrm{val}}(\theta_\lambda)

Hyperparameters are outside the ordinary parameter updates of the specified training run. They need not be selected manually: a separate search or optimization procedure may choose them. The distinction concerns the level of optimization being described. [2]

5. Regularization as an explicit example

For batch targets y∈Rny\in\mathbb R^n, a regularized linear-regression objective is:

J(w,b)=12n∥Xw+b1n−y∥22+ρ2∥w∥22,ρ≥0J(w,b)=\frac{1}{2n}\|Xw+b\mathbf1_n-y\|_2^2 +\frac{\rho}{2}\|w\|_2^2,\qquad \rho\ge0

The first term measures mean squared residual with a factor of 1/21/2. The second penalizes weight magnitude; the bias is not penalized in this definition. The model parameters are w,bw,b, while ρ\rho controls the relative strength of the penalty.

Differentiating the penalty gives ρw\rho w, which is added to the unregularized weight gradient. Thus a hyperparameter can directly affect parameter updates without itself being updated by that same rule. A complete specification must identify the adjustable variables, fixed quantities and selection procedure; labels alone are insufficient.

References