AI coursesMain page
Back to the library

Generalization, overfitting and model evaluation

Understand what a training fit says about new data.

On this page
  1. Population The corresponding empirical risk
  2. Non-uniqueness of a training fit
  3. Training, validation and test sets
  4. Accuracy and sampling uncertainty
  5. Overfitting, metric choice and distribution shift
  6. References

The purpose of fitting is to make useful predictions on new samples. A model can match every training point while behaving very differently between or beyond those points. Generalization analysis asks which of those possible extensions is supported by the available data, and held-out evaluation measures how the selected extension performs.

Generalization concerns performance on samples beyond those used for fitting. Its meaning depends on the target distribution and evaluation loss. A low training error alone does not establish generalization, because parameter selection has already used the training observations.

1. Population The corresponding empirical risk

Let PP denote the intended joint distribution of inputs and targets. For fixed parameters θ\theta, population risk is:

RP(θ)=E(x,y)∼P[ℓ(fθ(x),y)]R_P(\theta)=\mathbb E_{(x,y)\sim P} [\ell(f_\theta(x),y)]

The corresponding empirical risk on nn observed samples is:

R^n(θ)=1n∑i=1nℓ(fθ(xi),yi)\widehat R_n(\theta)=\frac1n\sum_{i=1}^{n} \ell(f_\theta(x_i),y_i)

Population risk is generally unknown. A held-out average estimates it under appropriate sampling and independence conditions: the evaluation samples must represent the intended distribution and remain independent of the fitted procedure. This independence separates evaluation from reuse of the observations that determined the parameters. [1]

2. Non-uniqueness of a training fit

Consider training pairs (−1,1)(-1,1), (0,0)(0,0) and (1,1)(1,1), and the function family:

fc(x)=x2+c x(x2−1),c∈Rf_c(x)=x^2+c\,x(x^2-1),\qquad c\in\mathbb R

At all three training inputs, x(x2−1)=0x(x^2-1)=0, so every value of cc gives zero training squared error. If the target relation outside those points is explicitly specified as y=x2y=x^2, then at x=2x=2:

fc(2)=4+6c,(fc(2)−4)2=36c2f_c(2)=4+6c,\qquad (f_c(2)-4)^2=36c^2

The same training error is compatible with arbitrarily different error at a new point. The training observations alone leave the coefficient cc undetermined; selecting a useful extension requires an additional criterion or evidence from other samples.

3. Training, validation and test sets

Training data fit parameters; validation data select models or configurations; test data evaluate the selected procedure. If θλ\theta_\lambda results from training with configuration λ\lambda, validation selection can be written as:

λ∗∈arg⁡min⁡λ∈ΛR^val(θλ)\lambda^*\in\arg\min_{\lambda\in\Lambda} \widehat R_{\mathrm{val}}(\theta_\lambda)

The held-out test set then estimates the selected procedure's risk. Repeatedly using its results to guide unrestricted tuning compromises its role as independent evaluation. [2]

Data-dependent preprocessing must respect the same boundary. Normalization statistics and feature selection should be fitted using the appropriate training portion, then applied to held-out samples. Selecting features from test labels before fitting the predictor already introduces leakage. [3]

The sampling unit also matters. Records from the same source or nearby positions in a correlated sequence may not be independent. Group-based or temporal splits can be necessary when random row-level splitting does not match the intended evaluation setting.

4. Accuracy and sampling uncertainty

For a fixed classifier and independent identically distributed test samples, let Zi=1[y^i=yi]Z_i=\mathbf1[\hat y_i=y_i] and p=Pr⁡(y^=y)p=\Pr(\hat y=y). Then:

p^=1m∑i=1mZi,E[p^]=p,Var⁡(p^)=p(1−p)m\hat p=\frac1m\sum_{i=1}^{m}Z_i, \qquad \mathbb E[\hat p]=p, \qquad \operatorname{Var}(\hat p)=\frac{p(1-p)}{m}

Since Zi2=ZiZ_i^2=Z_i, each indicator has variance p−p2p-p^2. Independence removes cross-covariance terms, giving the variance of the average shown above. Replacing pp by p^\hat p yields the plug-in standard-error estimate p^(1−p^)/m\sqrt{\hat p(1-\hat p)/m}, not an exact confidence interval.

For m=100m=100 and p^=0.9\hat p=0.9, the estimated standard error is 0.030.03. The formula does not apply unchanged to strongly correlated outcomes or a classifier repeatedly selected against the same test set.

5. Overfitting, metric choice and distribution shift

Overfitting occurs when adaptation to sample-specific variation does not carry over to the intended distribution. Further fitting may lower training loss while worsening held-out loss under comparable conditions. A nonzero gap alone does not identify its cause; distribution differences, leakage, sample size and label quality also affect comparisons.

Metrics impose another limitation. On a binary dataset with 90 positive and 10 negative labels, always predicting positive gives 90% accuracy but 0% recall for the negative class. Evaluation therefore requires the class structure and relevant error costs, not accuracy alone.

Finally, deployment under a distribution QQ concerns RQ(θ)R_Q(\theta) rather than RP(θ)R_P(\theta). Even an unbiased estimate on PP does not guarantee performance on QQ. Evaluation claims must state which distribution, loss and sampling conditions they concern. [1]

References