Skip to content

Temporal Conv Nets are basically normal CNNs, however they have 1) Dilated Convolutional Layers, 2) Causal Convolutions

Question list

  • what is the shape of the temporal kernel?
  • Are the temporal layers analyzed separately from the spatial layers? or do we use a 3D filter to analyze spatiotemporal relationships?
    • 3D kernels arent great for my project because there's no continuity/co-registration between images
  • How are filter values actually learned?. How does backprop work with CNNs?

Foundational Concepts

Receptive Field

The receptive field is how "expansive our eyesight is". How far in the past we see. how many pixels' signals reach a single hidden node.

Conv 1D for time series

We can use Conv1D to analyze time series (a continuous stream of a single value over time) by some simple 1D kernel like [1 1 1] (which provides a moving window average), where each node is a single value.

Analogously, if we have a hidden state at each timepoint, we can still apply a kernel (what shape? 2D?) to obtain information about hidden states. However, standard convolution also considers values from the future! This is cheating - we can't describe a past state in terms of a future state! So we use Causal Convolutions to censor the future values

Causal Convolutions

When analyzing a given timepoint t, we only expose it to data from timepoints < t

This is achieved by zero-padding at the start of the sequence, rather than padding at both ends as is tyically done in CNNs

Dilated Convolution

We could analyze one timepoint in the past easily with filters, but if we want any sense of long-range dependencies we would need an enormous filter spanning all the timepoints. This already amounts to an enormous number of parameters in 1D timeseries space - imagine the cost when it's 2D!

The goal of dilated convolution is to look at the current timepoint, \(t\), and one single other timepoint in the past, \(t - d\) . \(d\) is the dilation factor, defining the distance back in time we look. This allows us to expand our receptive field linearly with the number of parameters we must learn.

In order to ensure we get signal from all timepoints, we typically use a pyramidal dilation scheme in which later layers use larger and larger dilation factors. (e.g convolutional layer \(l\) uses dilation factor = \(2^l\)), which seems analogous to how in Transformers we use more and more encoding blocks to see more of the input. D = {1, 2, 4, 8, ...}

For a 1D signal, with L total layers and dilation scheme \(d = 2^l\), the receptive field is \(\(R = 2^L(k-1)\)\)

Properties of TCNs

  • highly parallelized like transformers, with the added benefit of not being as data demanding since we're learning filters!
    • I think i did notice that TCNs are faster to train
  • gradient flow is unlike RNNs where you need to unfold such that the output of one hidden state depends on past hidden states. Here, we know the computation of layer x depends on the values of \(x\) and \(x-d\), so we can compute gradients in parallel
  • no vanishing gradient since we continue to look at early inputs in the sequence even at the final layer! We look at all inputs in the sequence at once so all elements are equally weighted.
  • The receptive field is fixed a priori according to our architecture's dilation factors. Unlike #transformers, #TCNs cannot view inputs all at once irrespective of sequence size.
  • greater control over receptive field. We can modify 1) kernel size 2) network depth 3) dilation factor to increase receptive field.

Takeaways

  • positional padding is inappropriate for TCNs, as the signal would simply be lost given the fixed receptive field
  • inherently, solutions which examine elements in a sequence in parallel like #transformers and #TCNs do not struggle with vanishing gradient problem.