machinelearning:: Loss functions are nonconvex, so local minima != global minima. But local minima actually arent a huge concern in most cases. Local minima tend to be similar in cost to the global minimum. If we think we've converged but still are getting bad results, the culprit could be a bad local minimum. Deep Learning Book
Summary¶
- To test whether we're in a local minumim, we want to see the size of our gradient. Given that the gradient at every layer is a vector, we can take the norm of the gradient to see the size of change. If we're not changing much, then we're probably in a local minimum.