Sources for representative benchmarks
- Dominant language
- No language data
- Stars
- 25
- Forks
- 3
- PR merge metrics
- No merged PRs in 30d
Description
Over the years we've had good results with profiling specific (and _reproducible_) benchmarks from the larger pydata community and using the results to make further improvements to dask. Good workloads have come from community blogposts, performance issues from the pangeo community, etc... It would be good to periodically collect good large benchmarks like this to perform and use for motivating and tracking larger performance changes. Random datasets can be great for tests, but aren't always the most representative of real world problems.
Opening this here as a place to solicit/collect links to notebooks/blogposts/etc... that provide reproducible examples we can use to profile and improve dask over time.
Contributor guide
Assessment
This issue has not been assessed yet.