Duration 15:14
16+
Play
Video

Mesh-TensorFlow: Model Parallelism for Supercomputers

Noam Shazeer
Software Engineer at Google
  • Video
  • Video
TensorFlow Dev Summit 2019
March 7 2019, Sunnyvale, CA, United States
TensorFlow Dev Summit 2019
Video
Mesh-TensorFlow: Model Parallelism for Supercomputers
Purchased
In cart
Free
Free
Free
Free
Free
Free
Add to favorites
6.52 K
I like 0
I dislike 0
Purchased
In cart
Free
Free
Free
Free
Free
Free
  • Description
  • Discussion

About speaker

About the talk

Batch-splitting (data-parallelism) is the dominant distributed Deep Neural Network (DNN) training strategy, due to its universal applicability and its amenability to Single-Program-Multiple-Data (SPMD) programming. However, batch-splitting suffers from problems including the inability to train very large models (due to memory constraints), high latency, and inefficiency at small batch sizes. All of these can be solved by more general distribution strategies (model-parallelism). Unfortunately, efficient model-parallel algorithms tend to be complicated to discover, describe, and to implement, particularly on large clusters. We introduce Mesh-TensorFlow, a language for specifying a general class of distributed tensor computations. Where data-parallelism can be viewed as splitting tensors and operations along the "batch" dimension, in Mesh-TensorFlow, the user can specify any tensor-dimensions to be split across any dimensions of a multi-dimensional mesh of processors. A Mesh-TensorFlow graph compiles into a SPMD program consisting of parallel operations coupled with collective communication primitives such as Allreduce. We use Mesh-TensorFlow to implement an efficient data-parallel, model-parallel version of the Transformer sequence-to-sequence model. Using TPU meshes of up to 512 cores, we train Transformer models with up to 5 billion parameters, surpassing state of the art results on WMT'14 English-to-French translation task and the one-billion-word language modeling benchmark. Mesh-Tensorflow is available at https://github.com/tensorflow/mesh .

Speaker: Noam Shazeer, Google

Share

Cackle comments for the website

Buy this talk

Access to the talk «Mesh-TensorFlow: Model Parallelism for Supercomputers»
Purchased
In cart
Free
Free
Free
Free
Free
Free

Video

Get access to all videos “TensorFlow Dev Summit 2019”
Purchased
In cart
Free
Free
Free
Free
Free
Free
Ticket

Similar talks

Nathan Silberman
Lead Deep Learning Scientist at Butterfly Network
Purchased
In cart
Free
Free
Free
Free
Free
Free
Wei Lin
Senior Director at Alibaba
Purchased
In cart
Free
Free
Free
Free
Free
Free
Krzysztof Ostrowski
Research Scientist at Google
Purchased
In cart
Free
Free
Free
Free
Free
Free

Buy this video

Video

Access to the talk 'Mesh-TensorFlow: Model Parallelism for Supercomputers'
Purchased
In cart
Free
Free
Free
Free
Free
Free

Conference Cast

With ConferenceCast.tv you get access to our library of the world's best conference talks.

Conference Cast
173 conferences
7344 speakers
2377 hours of content