dask / dask/distributed

Learn heterogeneous bandwidths

Open
#2,743 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
1.7k
Forks
778
Avg merge
2h 50m
Merged PRs (30d)
3

Description

In order to make good scheduling decisions the scheduler often has to make an estimate for how long transfers will take. Currently, it learns a uniform exponentially weighted moving average based on what the workers observe.

However, this assumption of uniformity breaks down in a few cases:

1. Different types often incur different serialization costs (which we bundle into bandwidth here)
2. Different types may also move over different transports, as with GPU data and NVLink
3. Different workers may be closer or farther away from each other. For example they may be on the same node, in the same rack, or in the same data center
4. Very small frames often have some baseline cost

Learning a model that estimates the total transit time of a piece of data would be useful, but it may also be somewhat tricky. There is a balance to be struck between generalizing across the cluster and data types and learning heterogeneity that may exist.

Also, this needs to be fairly lightweight on the scheduler.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.