Conceptual

Gradient Descent

Iteratively updates parameters by stepping opposite the loss gradient, with a learning rate controlling step size. Stochastic minibatch variants trade exact gradients for far cheaper, noisier steps.