2 articles
What a neuron actually computes, why the squashing function is the part that matters, and how backpropagation extracts every partial derivative in a single backward sweep.
How repeatedly stepping in the direction of steepest descent minimizes functions in millions of dimensions — and trains essentially every modern neural network.