AI, machine learning and deep learning
Separate the task, the model and the way it learns.
On this page
A learning system must connect three decisions: what to predict, how to represent the input, and how to improve its predictions from examples. Separating these decisions makes the relationship between AI, machine learning and deep learning concrete. The following formulation starts with a prediction task and shows which parts become trainable.
Artificial intelligence (AI), machine learning and deep learning refer to different levels of a technical framework. AI is the broader research field; machine learning comprises methods that infer useful structure from data; deep learning is a branch of machine learning based on multilayer neural networks. A task, a model family and an optimization method are separate concepts: classification specifies an output requirement, a neural network specifies a function family, and gradient descent specifies a parameter-update procedure. [1]
1. Learning as an optimization problem
A supervised dataset consists of input–target pairs:
Here is the number of samples and is the input dimension. A model produces the prediction , with adjustable parameters collected in . Given a task-specific loss , the empirical risk is:
Training seeks parameters that reduce this objective, possibly with an additional regularization term. Empirical risk measures fit to the observed dataset; generalization evaluates the resulting model on new samples. [2]
2. Classification, regression and generation
The distinction between tasks depends on the meaning and structure of the target. [2]
| Task | Target | Example formulation |
|---|---|---|
| Regression | A real-valued quantity | , with squared loss |
| Classification | A category | Class probabilities |
| Sequence generation | A token sequence | Conditional distributions over successive tokens |
For regression, a common loss is . For classification, the negative log-likelihood of the target category is . For an autoregressive sequence model, the objective sums negative log-probabilities over target tokens. Although each implementation operates on numbers, these objectives encode different prediction requirements.
Training setup is another distinction. Supervised learning uses supplied targets, whereas self-supervised learning constructs targets from the data itself, such as the next token of an observed sequence. Generation is therefore compatible with self-supervised training; output type and target-construction method are not competing categories.
3. Fixed and learned representations
A model based on fixed features can be written as:
The representation is specified in advance and the prediction parameters are fitted. When the representation is trainable, the model becomes:
Optimization can update both and . A fixed-feature classifier already learns ; representation learning adds to the optimized variables. This comparison isolates one change in the learning pipeline.
A multilayer representation has the form:
Here counts successive transformations. The matrices and biases have compatible shapes, and denotes an activation function. Depth describes function composition. Its effect on expressive capacity depends on the transformations involved: a stack of affine maps without intervening nonlinearities remains affine. [1]
4. A two-sample regression example
Consider the constructed samples and with . Their average squared error is:
Thus and , with the unique minimizer . The next question is whether the fitted relation also describes new samples, which motivates the evaluation procedures in article 5.
5. Model and application system
A trained function is typically one component of an application. Preprocessing, retrieval, decoding and output handling can alter system behavior without changing its stored parameters. Consequently, a statement about the model architecture does not fully specify how an application produces an answer.
The relevant analytical objects are the data, representation, model, objective, optimization procedure and evaluation distribution. Keeping these objects distinct provides a consistent basis for comparing traditional learning methods, deep networks and language models.