Skip to content

machinelearning:: depthwise separable convolution operates on space and depth in two separate steps to be more computationally efficient.

Summary

Ideas

  • Depthwise separable convolution: operates on space and depth in two steps.
    • depthwise convolution: applies a convolutional kernel to each input channel
    • pointwise convolution: applies a \(1 \times 1\) convolution to combine the channel-wise outputs of depthwise convolution
      • this ends up being more computationally efficient than learning a 3D kernel
      • i could possibly apply the same concept along the time dimension to characterize the change across time. I'm not really sure where the gating mechanism comes in.

can i apply RNN attention gating to MHA?

Abstract

A multi-scale gated multi-head attention depthwise separable CNN model for recognizing COVID-19 | Scientific Reports (nature.com)