Skip to content

Gated Recurrent Unit

Similar in concept to LSTM: We use an update gate, reset gate to manage memory content.

Update Gate

The update gate controls how much of the past information needs to be passed along to the future. The model could decide to copy all information from the past and eliminate the vanishing gradient problem, or hold onto the most salient information. $$ z_t = \sigma (W^zx_t + U^zh_{t-1}) $$ where \(x_t\) is the input t, and \(h_{t-1}\) is the hidden state which accumulates all information from t-1 units. W and U are learned matrices to transform \(x_t\) and \(h_{t-1}\) into a shared vector space, whereby their scalar values are weighted by their importance to the new state. Then we apply a sigmoid to squash results between 0 and 1.

remaining questions:

how does \(h_{t-1}\) hold onto all past information? why sigmoid?

Reset Gate

Decide how much past information to forget $$ r_t = \sigma (W^rx_t + U^rh_{t-1}) $$ This is the same formula as the update gate, just that the learned matrices are different.

Current memory content

Forgetting

A two-step process First we obtain a representation of the input, contextualized with history after forget-gate is applied.

$$h_t' = tanh(Wx_t + r_t \otimes Uh_{t-1}) $$ 1) \(Wx_t\) for a linear transform of the input 2) The hadamard (element-wise) product between the reset gate \(r_t\) and \(Uh_{t-1}\) 1) recall that \(r_t\) varies with time depending on Reset Gate , the goal is to learn values at each time point indicating the importance of each past feature to the final classification 2) sum up the results from 1 and 2, assuming \(W\) and \(U\) brought them into the same latent space, and apply a nonlinearity.

Step 2 Remembering

$$ h_t = z_t \odot h_{t-1} + (1-z_t)\odot h'_t $$ 1) elementwise multiplication to the update gate \(z_t\) and past hidden state \(h_{t-1}\) this first term tracks what we should remember from the past history 1) element-wise multiplication to \((1-z_t)\) and the intermediate state this second term \(h'_t\) tracks the current input, contextualized by some of the past history

Generalizing to transformers

X Gated Attention for Recency Bias uses this same gating mechanism to impose a recency bias for transformer self-attention. - I could probably use the same idea of applying a gating mechanism to self-attention in order to promote self-attention between entries close in time X Gated Attention for Recency Bias#Ideas - could we replace the positional encoding step with our time encoding instead? - there are various desirable traits of positional embeddings, covered in positional embeddings - could we learn a time embedding through a function f(t)? and take the hadamard product instead of