apache / apache/datafusion-comet

Enable the split-operator plan and native Iceberg writes by default

Open
#5,644 0 comments 0 reactions 0 assignees View on GitHub
area:ci area:Iceberg area:writer enhancement requires-triage
Dominant language
Scala
Stars
1.3k
Forks
373
Avg merge
2d 6h
Merged PRs (30d)
190

Description

### What is the problem the feature request solves?

Native Iceberg writes ship behind two flags that both default to `false`: `spark.comet.write.iceberg.splitOperator.enabled` (the writer/committer split from #4658) and `spark.comet.iceberg.write.enabled` (the iceberg-rust data-file writer from #5361). Nothing in CI runs with either flag on, so the feature is only exercised by the suites that set the flags themselves, and there is no agreed definition of what has to be true before the defaults flip.

#5259 catalogued what breaks when the split plan is enabled by default (8 jobs, 4 buckets). That is the first step, but it does not cover the native writer flag or say what else is required.

### Describe the potential solution

Graduation criteria, each of which should be a linked issue or a checked item before the defaults change:

- [ ] #5259: the four failure buckets when the split plan is on by default are fixed
- [ ] A CI job in `pr_build_linux.yml` runs the Iceberg suites with both flags on (or the suites default them on), so a regression on the native path fails a PR rather than a manual run
- [ ] The two correctness bugs on the native path are fixed: #5636 (partition-path escaping) and #5637 (GCS configuration)
- [ ] Failure handling matches iceberg-java: #5618 (task-attempt cleanup) and #5277 (orphans on commit failure), with failure-injection tests
- [ ] Manifest metrics parity is asserted per data type across all four Iceberg versions in the build matrix (1.5.2, 1.8.1, 1.10.0, 1.11.0), not only for the hand-picked cases in `CometIcebergWriteActionSuite`
- [ ] The default hash-distribution write is native end to end (#5635)
- [ ] The Spark SQL / Iceberg test diffs under `dev/diffs` pass with the flags on
- [ ] A write benchmark exists and shows no regression against iceberg-java for unpartitioned, clustered, and fanout writes
- [ ] The eligibility restrictions each have a keep-or-lift decision, so the fallback surface is intentional

Suggested rollout: flip `splitOperator.enabled` first (it changes plan shape but still writes through iceberg-java, so the blast radius is planning only), then `iceberg.write.enabled` one release later.

### Additional context

Part of the native Iceberg writes epic, #5649. Related: #4658, #5298, #5361.

Contributor guide

Open the contributing guide

Research direction

Start with pr_build_linux.yml and the Iceberg suites, including CometIcebergWriteActionSuite, then review the listed issues and the Spark SQL/Iceberg diffs under dev/diffs. Check the four-version build matrix and the requested failure-injection and benchmark coverage. Done means every graduation criterion is resolved or checked before either default changes.

Written by the indexing model from the issue text.

Assessment

Tech stack
rust, scala
Domain
ci-cd, data-engineering, databases
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.