Skip to content

machinelearning:: Tips for how to add parallelism into our training (data parallelism and pipeline parallelism)

Summary

  • Data parallellism: we feed in different batches of data into different GPUs
    • challenge is how do we sync the calculated gradients to ensure model parameters are updated appropriately?
      • tend to calculate an average: but still message passing between GPU can slow down performance. There are some special synchronous techniques out there for this i think?
  • Pipeline Parallelism: sequential layers of the model are stored on different GPU: this way, no single GPU needs to store the entire model (this is probably beyond the scope of my relatively small models) - issue is then we have waiting time as we need one region of the model to finish in order for the next region to start - to address this, we can split batches up into minibatches, so that the later GPUs can receive information quicker -  Strategies for managing these differs. "GPipe has each worker process forward and backward passes consecutively and then aggregates gradients from multiple microbatches synchronously at the end. PipeDream instead schedules each worker to alternatively process forward and backward passes."
  • Tensor Parallelism: We can split computations up between GPUs. for example, the bottleneck with transformers is matrix multiplication of attention matrices and weight matrices. Mat Mult is essentially dot products between rows and columns. So we can use GPUs to perform dot products between different rows and columns of the same tensor.
    • this is a nice note of why it's important to understand where the bottlenecks in runtime are so that i can learn how to address it better.

Ideas

Important takeaway: I should think about bottlenecks to my own models. For example, the sequential models i've written have a high fixed I/O cost when creating the initial csv containing all the data. Then when we load the features, rather than read them one at a time (which would probably be quite slow as another I/O operation), we load them all into memory at once. But this comes at the cost of more memory being occupied since we need all the features at once to be loaded up in memory. And as we learned in operating systems, having low RAM leads to slow down because we need to continually put things into solid state drive instead of caching.

Abstract

Techniques for Training Large Neural Networks (openai.com)