NBytes estimates on pandas dataframes with strings are overly pessimistic
Open
- Dominant language
- Python
- Stars
- 1.7k
- Forks
- 778
- Avg merge
- 2h 50m
- Merged PRs (30d)
- 3
Description
This results in Dask swapping to disk needlessly. Now that we're better at checking process-level memory we might consider relaxing our spill-to-disk mechanism based on estimated byte costs on a per-data-element basis.
Contributor guide
Assessment
This issue has not been assessed yet.