machinelearning:: replacing every convolution layer w/ MHA leads to improved performance on resnet. MHA is more effective on later layers Summary¶ Ideas¶ Abstract¶ Stand-Alone Self-Attention in Vision Models (neurips.cc)