dask / dask/community

Sources for representative benchmarks

Open
#192 2 comments 1 reaction 0 assignees View on GitHub
Dominant language
No language data
Stars
25
Forks
3
PR merge metrics
No merged PRs in 30d

Description

Over the years we've had good results with profiling specific (and _reproducible_) benchmarks from the larger pydata community and using the results to make further improvements to dask. Good workloads have come from community blogposts, performance issues from the pangeo community, etc... It would be good to periodically collect good large benchmarks like this to perform and use for motivating and tracking larger performance changes. Random datasets can be great for tests, but aren't always the most representative of real world problems.

Opening this here as a place to solicit/collect links to notebooks/blogposts/etc... that provide reproducible examples we can use to profile and improve dask over time.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.