apache / apache/beam

Allow non-deferred column operations on categorical columns

Open
#20,958 2 comments 0 reactions 0 assignees View on GitHub
core dataframe improvement P3 python
Dominant language
Java
Stars
8.7k
Forks
4.7k
Avg merge
1d 20h
Merged PRs (30d)
196

Description

There are several operations that we currently disallow because they produce a variable set of columns in the output based on the data (non-deferred-columns). However, for some dtypes (categorical, boolean) we can easily enumerate all the possible values that will be seen at execution time, so we can predict the columns that will be seen.

Note we still can't implement these operations 100% correctly, as pandas will typically only create columns for the values that are __observed__, while we'd have to create a column for every possible value.

We should allow these operations in these special cases.

Operations in this category:
- DataFrame.unstack, Series.unstack (can work if unstacked level is a categorical or boolean column)
- Series.str.get_dummies
- Series.str.split
- Series.str.rsplit
- DataFrame.pivot
- DataFrame.pivot_table

Imported from Jira [BEAM-12169](https://issues.apache.org/jira/browse/BEAM-12169). Original Jira may contain additional context.
Reported by: bhulette.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.