apache / apache/beam

DataFrame API: Consider allowing partitioning by column in addition to Index

Open
#20,857 0 comments 0 reactions 0 assignees View on GitHub
core dataframe dsl improvement P3 python
Dominant language
Java
Stars
8.7k
Forks
4.7k
Avg merge
1d 20h
Merged PRs (30d)
196

Description

For some DataFrame use-cases it may be beneficial to partition a dataset across the columns as well as across the index.

One example might be computing a correlation in a DataFrame with a very large number of columns. It would be beneficial to be able to perform pairwise column correlations on separate workers.

Imported from Jira [BEAM-12132](https://issues.apache.org/jira/browse/BEAM-12132). Original Jira may contain additional context.
Reported by: bhulette.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.