Skip to content

machinelearning:: The #LotteryTicketHypothesis posits that a small subset of large, randomly initialized networks are responsible for the majority of the performance. #Pruning of the network to isolate these winning subnetworks can lead to improved generalization accuracy in shorter training time.

Summary

  • For #pruning the network they simply removing some percentage of the lowest-weight parameters (with some advanced strategies related to the learning-rate as well) within each layer.
    • then they reset the remaining connections to their original initializations and repeat (retraining from scratch but on a pruned network with the "good" initializations sticking)
  • Pruned models tend to train faster and generalize better to the test set.

The Lottery Ticket Conjecture. Returning to our motivating question, we extend our hypothesis into an untested conjecture that SGD seeks out and trains a subset of well-initialized weights. Dense, randomly-initialized networks are easier to train than the sparse networks that result from pruning because there are more possible subnetworks from which training might recover a winning ticket.

Ideas

- Given that they can prune the network by removing the lowest weight parameters, would it be helpful if we visualized the weight parameters after the model is done training? Would this assist in interpretability, and helping us understand the utility of our model?

Abstract

Neural network pruning techniques can reduce the parameter counts of trained networks by over 90%, decreasing storage requirements and improving computational performance of inference without compromising accuracy. However, contemporary experience is that the sparse architectures produced by pruning are difficult to train from the start, which would similarly improve training performance. We find that a standard pruning technique naturally uncovers subnetworks whose initializations made them capable of training effectively. Based on these results, we articulate the lottery ticket hypothesis: dense, randomly-initialized, feed-forward networks contain subnetworks (winning tickets) that—when trained in isolation— reach test accuracy comparable to the original network in a similar number of iterations. The winning tickets we find have won the initialization lottery: their connections have initial weights that make training particularly effective. We present an algorithm to identify winning tickets and a series of experiments that support the lottery ticket hypothesis and the importance of these fortuitous initializations. We consistently find winning tickets that are less than 10-20% of the size of several fully-connected and convolutional feed-forward architectures for MNIST and CIFAR10. Above this size, the winning tickets that we find learn faster than the original network and reach higher test accuracy.

The #LotteryTicketHypothesis . A randomly-initialized, dense neural network contains a subnetwork that is initialized such that—when trained in isolation—it can match the test accuracy of the original network after training for at most the same number of iterations. Lottery Ticket Hypothesis (arxiv.org)