dask / dask/distributed

NBytes estimates on pandas dataframes with strings are overly pessimistic

Open
#1,787 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
1.7k
Forks
778
Avg merge
2h 50m
Merged PRs (30d)
3

Description

This results in Dask swapping to disk needlessly. Now that we're better at checking process-level memory we might consider relaxing our spill-to-disk mechanism based on estimated byte costs on a per-data-element basis.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.