apache / apache/beam

Support PARTITIONED BY on Beam's SQL DDL

Open
#20,885 0 comments 0 reactions 0 assignees View on GitHub
dsl new feature P3 sql
Dominant language
Java
Stars
8.7k
Forks
4.7k
Avg merge
1d 20h
Merged PRs (30d)
196

Description

Partitioning by columns is a common optimization technique used in Hive and Spark to optimize queries performance.

Beam should support this feature to allow users that already have data stored following a partitioning schema to read and query it with Beam SQL.

Imported from Jira [BEAM-12315](https://issues.apache.org/jira/browse/BEAM-12315). Original Jira may contain additional context.
Reported by: iemejia.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.