apache / apache/paimon-cpp

[Feature] Maintain source-backed primary-key BTree indexes during compaction

Ouverte
#291 0 commentaires 0 réactions 0 personnes assignées Voir sur GitHub
Langage dominant
C++
Étoiles
65
Forks
25
Merge moyen
2 j 12 h
PR mergées (30 j)
80

Description

### Search before asking

- [x] I searched in the issues and found nothing similar.

### Motivation

#192 and #194 added the source-backed primary-key BTree read path. Paimon C++ writers still need the corresponding maintenance path: after compaction changes the active source files of a data level, a missing or stale payload leaves that level uncovered and queries fall back to normal file scans.

Paimon C++ should maintain these payloads during fixed-bucket primary-key writes and compaction, using the existing Java-compatible source metadata, BTree payload format, and index manifests.

### Solution

Add the source-backed primary-key BTree maintenance lifecycle for fixed-bucket primary-key tables:

1. Validate the Java-equivalent table and index prerequisites.
2. Restore committed source-backed payload metadata into bucket writers without mixing Data Evolution payloads.
3. Build one payload per indexed field and positive data level from physical source rows, then commit matching index additions and deletions in the same snapshot as the data changes.
4. Reconcile missing, stale, duplicate, replaced, removed-definition, and empty-level payloads during compaction.
5. Isolate build failures to the affected field and level so reads safely fall back to normal scans and a later maintenance attempt can rebuild the payload.
6. Retain live index files during snapshot expiration and orphan cleanup, including tag and branch safety and external-file deletion retries.

Reuse the existing storage formats and internal reader, writer, sort-buffer, path, manifest, and commit abstractions. Do not introduce a new index family or storage protocol.

The implementation is in #245. It keeps maintenance synchronous; Java asynchronous scheduling, manual rebuild actions, realtime writers, and postpone-bucket writers remain outside this scope.

### Anything else?

This is a maintenance-path follow-up to the read-path work in #192 and #194. It ports an existing Java capability, so no separate PIP is proposed.

### Are you willing to submit a PR?

- [x] I'm willing to submit a PR!

Guide de contribution

Ouvrir le guide de contribution

Piste de recherche

Commencez par lire l’implémentation référencée dans #245 ainsi que le chemin de lecture de la clé primaire fondé sur le code source de #192 et #194. Suivez les formats de stockage existants et les abstractions de reader et writer internes, de manifest, de path, de sort-buffer et de commit. Le travail est considéré comme terminé lorsque les écritures en fixed-bucket et la compaction préservent des payloads par champ et de niveau positif, tandis que le nettoyage et les builds échoués préservent un fallback de scan sûr.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
cpp
Domaine
databases
Type d'issue
Fonctionnalité
Difficulté
5/5
Temps estimé
Plus d'une semaine
Activité
À l'abandon
Clarté
Plutôt claire
Accessibilité débutants
25/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.