machinelearning:: depthwise separable convolution operates on space and depth in two separate steps to be more computationally efficient.
Summary¶
Ideas¶
- Depthwise separable convolution: operates on space and depth in two steps.
- depthwise convolution: applies a convolutional kernel to each input channel
- pointwise convolution: applies a \(1 \times 1\) convolution to combine the channel-wise outputs of depthwise convolution
- this ends up being more computationally efficient than learning a 3D kernel
- i could possibly apply the same concept along the time dimension to characterize the change across time. I'm not really sure where the gating mechanism comes in.
can i apply RNN attention gating to MHA?