apache / apache/iceberg

Implement File Format API Technology Compatibility Kit

Open
#15,415 2 comments 2 reactions 0 assignees View on GitHub
improvement
Dominant language
Java
Stars
9.2k
Forks
3.5k
Avg merge
2d 11h
Merged PRs (30d)
132

Description

### Feature Request / Improvement

We need to have a set of tests for the new File Format API (#12774)

We need test for the features collected [here](https://docs.google.com/document/d/1sF_d4tFxJsZWsZFCyCL9ZE7YuI7-P3VrzMLIrrTIxds/edit?tab=t.0#heading=h.hrt6siqyob8m).

We can start from `TestGenericFormatModels`. We need tests writes with Generic data, and then test reading them with Generic/Spark/Spark vector/Flink/Arrow, then writes with Generic/Spark/Flink and reads with Generic.

The test suite should be easy to extend for new File Formats

### Query engine

None

### Willingness to contribute

- [ ] I can contribute this improvement/feature independently
- [ ] I would be willing to contribute this improvement/feature with guidance from the Iceberg community
- [ ] I cannot contribute this improvement/feature at this time

Contributor guide

Open the contributing guide

Research direction

Start with TestGenericFormatModels and the feature list in the linked document. Build the compatibility tests around Generic, Spark, Spark vector, Flink, and Arrow read/write combinations, with coverage for adding new file formats. Done means the requested combinations are tested and the suite is easy to extend.

Written by the indexing model from the issue text.

Assessment

Tech stack
java, spark
Domain
data-engineering, testing
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
52/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.