NVIDIA / NVIDIA/cudf

[FEA] Make `cudf::hash_partition` less delicate for large numbers of partitions

Open
#21,299 2 comments 0 reactions 0 assignees View on GitHub
feature request libcudf
Dominant language
C++
Stars
9.8k
Forks
1.1k
Avg merge
3d 6m
Merged PRs (30d)
278

Description

**Is your feature request related to a problem? Please describe.**

`cudf::hash_partition` is very delicate when the number of requested partitions is large enough that we don't dispatch to "optimized" kernels (`num_partitions > 1024`).

In the following ways:

1. The approach of tracking assignment of rows to partitions allocates a vector that turns out to be sparse in this case (so it can be much larger than the number of input rows, leading to out of memory).
2. Even if we get through that, the same sparse vector ends up with more than uint32::max values, and so we hit the usual 32bit offset thrust errors.

It would be great if we could lift these restrictions.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.