DataFrame API: Consider allowing partitioning by column in addition to Index
Open
core
dataframe
dsl
improvement
P3
python
- Dominant language
- Java
- Stars
- 8.7k
- Forks
- 4.7k
- Avg merge
- 1d 20h
- Merged PRs (30d)
- 196
Description
For some DataFrame use-cases it may be beneficial to partition a dataset across the columns as well as across the index.
One example might be computing a correlation in a DataFrame with a very large number of columns. It would be beneficial to be able to perform pairwise column correlations on separate workers.
Imported from Jira [BEAM-12132](https://issues.apache.org/jira/browse/BEAM-12132). Original Jira may contain additional context.
Reported by: bhulette.
Contributor guide
Assessment
This issue has not been assessed yet.