apache / apache/paimon-cpp

[Feature] Roadmap for Paimon C++ 0.4.0

オープン
#186 コメント 0 件 リアクション 2 件 担当者 0 名 GitHub で見る
enhancement
主要言語
C++
スター
65
フォーク
25
平均マージ
2日 12時間
マージ済み PR(30日)
80

説明

### Search before asking

- [x] I searched in the [issues](https://github.com/apache/paimon-cpp/issues) and found nothing similar.

### Motivation

Looking ahead, Paimon C++ will focus on several areas:

- Storage and retrieval optimizations for vertical workloads, including efficient MAP storage, richer data types such as VECTOR, and complete index support.
- Making newly written data queryable in real time through pluggable real-time writes and memory/disk union reads.
- Following up on important capabilities from the broader Paimon community, including changelog production, Format Table, and extensions to Data Evolution.

### Solution

#### 1. Storage and retrieval optimizations for vertical workloads

- [x] Support shared-shredding columnar storage for `MAP` based on [PIP-43](https://cwiki.apache.org/confluence/spaces/PAIMON/pages/430408347/PIP-43+Columnar+Storage+Optimization+for+MAP+Type+in+Paimon), including adaptive physical columns, field mappings, overflow storage, and end-to-end reads and writes.
- [ ] Support richer data types, starting with `VECTOR` based on [PIP-40](https://cwiki.apache.org/confluence/spaces/PAIMON/pages/399279132/PIP-40+Introduce+a+new+Vector+data+type), including schema representation, storage, reads, writes, and Data Evolution.
- [ ] Complete File Index support by adding index generation to the existing read path.
- [ ] Continue improving storage layout, predicate pushdown, column pruning, point lookups, and index-assisted retrieval for workload-specific scenarios.

#### 2. Making real-time data queryable

- [ ] Support pluggable real-time writes and memory/disk union reads based on [PIP-46](https://cwiki.apache.org/confluence/spaces/PAIMON/pages/444334302/PIP-46+Support+pluggable+real-time+writes+and+memory+disk+union+reads+for+Paimon+C), tracked by [#158](https://github.com/apache/paimon-cpp/issues/158).
- [ ] Make data held by an active writer queryable before it is committed into a snapshot.
- [ ] Guarantee consistent query results without missing or duplicate rows while writing, committing, and reclaiming memory segments concurrently.
- [ ] Support append tables first, followed by primary-key tables, deletion-vector mode, and additional production scenarios.

#### 3. Following up on Paimon community capabilities

- [ ] Support changelog production, including `input`, `lookup`, and `full-compaction` modes, tracked by [#174](https://github.com/apache/paimon-cpp/issues/174).
- [ ] Support Format Table reads and writes for Hive-style file directories, tracked by [#170](https://github.com/apache/paimon-cpp/issues/170).
- [ ] Extend Data Evolution to work with deletion vectors, tracked by [#169](https://github.com/apache/paimon-cpp/issues/169).
- [ ] Extend Data Evolution to primary-key tables and support compaction across evolved field groups, tracked by [#204](https://github.com/apache/paimon-cpp/issues/204).
- [ ] Continue following relevant Paimon features and make them available through native Paimon C++ APIs where appropriate.

Each roadmap item should be delivered through focused issues and reviewable pull requests, with unit tests, integration tests, documentation, and compatibility validation where applicable.

### Anything else?

_No response_

### Are you willing to submit a PR?

- [x] I'm willing to submit a PR!

コントリビューションガイド

コントリビューションガイドを開く

調査の方向性

リンクされている絞り込まれた issue から始めます。特に、リアルタイム書き込みと union 読み取りについては #158、changelog の生成については #174、Format Table については #170、Data Evolution については #169 または #204 を確認してください。該当する場合は参照先の PIP ドキュメントを読み、スコープを限定した roadmap 項目を 1 つ選びます。完了の条件は、関連する unit テストと integration テスト、ドキュメント、互換性検証を含む、焦点が明確でレビュー可能な変更です。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
cpp
領域
databases
issue の種類
機能追加
難易度
5/5
見積もり時間
1週間以上
活発さ
静か
明瞭さ
説明が足りない
初心者へのやさしさ
25/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。