apache / apache/beam

[Feature Request]: Add a Variant type to Beam Schemas

Open
#38,251 1 comment 0 reactions 0 assignees View on GitHub
java new feature P2 python
Dominant language
Java
Stars
8.7k
Forks
4.7k
Avg merge
1d 20h
Merged PRs (30d)
196

Description

### What would you like to happen?

PCollections of Beam Rows are required to have a fixed schema, making it hard to read or write records with varying logical schemas in the same PCollection.

We need a semi-structured type that can contain a nested level of variable columns. Transforms can choose to unwrap this type to reconstruct the variable columns accordingly.

The Variant type is becoming a good standard across different projects (see [Parquet](https://parquet.apache.org/docs/file-format/types/variantencoding/), [Spark](https://spark.apache.org/docs/latest/api/python/reference/pyspark.sql/api/pyspark.sql.types.VariantType.html), [Flink](https://nightlies.apache.org/flink/flink-docs-master/docs/sql/reference/data-types/#variant)). It could also be a good candidate for Beam.

Bonus if such a type is made portable so that other SDKs can use it too

### Issue Priority

Priority: 2 (default / most feature requests should be filed as P2)

### Issue Components

- [x] Component: Python SDK
- [x] Component: Java SDK
- [ ] Component: Go SDK
- [ ] Component: Typescript SDK
- [ ] Component: IO connector
- [ ] Component: Beam YAML
- [ ] Component: Beam examples
- [ ] Component: Beam playground
- [ ] Component: Beam katas
- [ ] Component: Website
- [ ] Component: Infrastructure
- [ ] Component: Spark Runner
- [ ] Component: Flink Runner
- [ ] Component: Samza Runner
- [ ] Component: Twister2 Runner
- [ ] Component: Hazelcast Jet Runner
- [ ] Component: Google Cloud Dataflow Runner

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.