Skip to content

machinelearning:: RNNs do be wildin RNNs in depth

Summary

  • I have a much better understanding of vanishing gradients for RNNs now!
    • With each timestep we must incorporate the previous state and the current input. So then the gradient becomes a sum of the previous state and the current timepoint. Take one more step, and now it must incorporate that past time point, the ex-current timepoint input, and the current timepoint. And so when we keep multiplyng fractions, we end up with smaller and smaller influences. leading to the issue of "vanishing gradients" in which the signal from initial timepoints is lost.
  • The book explains it much better, and if i want to become an expert in this field i need to do better.

Ideas

Abstract