machinelearning:: Models on the training set tend to exhibit double descent: as we get larger networks, training error will decrease, then increase, then decrease again. Summary¶ Ideas¶ Abstract¶ Deep Double Descent (openai.com)