apache / apache/iceberg

Core, Spark: Performant queries over (shredded) Variant data

Open
#16,172 1 comment 1 reaction 0 assignees View on GitHub
improvement
Dominant language
Java
Stars
9.2k
Forks
3.5k
Avg merge
2d 11h
Merged PRs (30d)
132

Description

Issue to group together everything needed for queries over Variant data to work well.

This is part of #10392

1. Auto generation of shredded fields.
2. Unmarshalling performance.
3. Rowgroup and file skipping based on shredded field stats.
4. Benchmarks to evaluate this

Iceberg query performance relies on spark to pass down variant_get() calls to the rowgroup filter, so the changes are interrelated. This stuff will have to target spark 4.2 only

[Proposal Document](https://docs.google.com/document/d/1IuhLRxw1rcPD_f4jgHuGe3SwFgy7Y5wgEGvLzf6311s/edit?usp=sharing)

## Iceberg

#14297 Spark: Support writing shredded variant in Iceberg-Spark
#15510 Parquet Rowgroup skipping for variant predicate
#15384 Api: Support variant extract and fix manifest bounds byte order
#15385 Spark: Support variant_get predicate pushdown for file skipping
#15628 Core, Spark: Add JMH benchmarks for Variants

+ skip files on iceberg stats, if possible.

## Spark

* [54598](https://github.com/apache/spark/pull/54598) Enable Parquet rowgroup skipping for variant filters
* [54394](https://github.com/apache/spark/pull/54394)
Support variant_get predicate for DSv2 filter pushdown

## Parquet: better unmarshalling

* [3452](https://github.com/apache/parquet-java/pull/3452) GH-3451. Add a JMH benchmark for variants
* [3481](https://github.com/apache/parquet-java/pull/3481) Optimizing Variant read path with lazy caching

### Query engine

Spark

### Willingness to contribute

- [ ] I can contribute this improvement/feature independently
- [x] I would be willing to contribute this improvement/feature with guidance from the Iceberg community
- [ ] I cannot contribute this improvement/feature at this time

Contributor guide

Open the contributing guide

Research direction

Start with the linked proposal document and the related Iceberg issues #14297, #15510, #15384, #15385, and #15628, along with the referenced Spark and Parquet pull requests. The scope covers shredded-field generation, unmarshalling, rowgroup and file skipping, and benchmarks, targeting Spark 4.2; done means these interrelated performance areas work together and are evaluated by benchmarks.

Written by the indexing model from the issue text.

Assessment

Tech stack
java, spark
Domain
data-engineering, databases, performance
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.