apache / apache/paimon-cpp

[Feature] Support changelog production

Aperta
#174 0 commenti 0 reazioni 1 assegnatario Assegnata a @lszskye Vedi su GitHub
enhancement
Lingua principale
C++
Stelle
65
Fork
25
Merge medio
2g 12h
PR unite (30g)
80

Descrizione

### Search before asking

- [x] I searched in the [issues](https://github.com/apache/paimon-cpp/issues) and found nothing similar.

### Motivation

Java Paimon supports producing changelog files for primary-key tables through the
`changelog-producer` option, including `input`, `lookup`, and `full-compaction`
modes.

Paimon C++ already contains part of the required infrastructure, such as:

- Parsing the `changelog-producer` option.
- Changelog file metadata in `DataIncrement` and `CompactIncrement`.
- Changelog manifest and snapshot commit support.
- Lookup changelog merge functions and changelog compaction rewriters.
- Streaming changelog scans and the `audit_log` and `binlog` system tables.

However, schema validation currently rejects every value except `none`, and the
writer does not produce changelog files. As a result, primary-key tables written
by Paimon C++ cannot provide the same changelog behavior as tables written by
Java Paimon.

Supporting Java-compatible changelog production would enable incremental
downstream consumption and improve interoperability between Java Paimon and
Paimon C++.

### Solution

Implement changelog production for Paimon C++ primary-key table writes, aligned
with Java Paimon semantics and storage format.

The implementation should cover the following producer modes:

- `input`
- Preserve input row kinds and double-write input changes to changelog files.
- Attach generated changelog files to the corresponding data increment.

- `lookup`
- Look up the previous value of each primary key when necessary.
- Produce the correct `INSERT`, `UPDATE_BEFORE`, `UPDATE_AFTER`, and `DELETE`
records.
- Generate changelog files while compacting level-0 data.

- `full-compaction`
- Produce changelog files during full compaction by comparing the compacted
result with the previous table state.

The generated changelog files should:

- Use the same file schema, row-kind encoding, metadata, and directory layout as
Java Paimon.
- Be committed through changelog manifests and referenced by the snapshot.
- Be readable through streaming scans, `audit_log`, and `binlog`.
- Work correctly with supported merge engines, sequence fields, partitions, and
bucket modes.
- Either support deletion vectors consistently with Java Paimon or reject
unsupported combinations with a clear validation error.
- Be included in commit metrics, expiration, rollback, and orphan-file cleanup.

The feature can be delivered incrementally, with separate pull requests for each
changelog producer mode.

### Anything else?

_No response_

### Are you willing to submit a PR?

- [x] I'm willing to submit a PR!

Guida per i contributori

Apri la guida per i contributori

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.