apache / apache/paimon-cpp

[Feature] Support richer data types, starting with VECTOR<T, N> based on PIP-40

オープン
#197 コメント 0 件 リアクション 0 件 担当者 1 名 @ChaomingZhangCN が担当を希望しています GitHub で見る
enhancement
主要言語
C++
スター
65
フォーク
25
平均マージ
2日 12時間
マージ済み PR(30日)
80

説明

### Search before asking

- [x] I searched in the [issues](https://github.com/apache/paimon-cpp/issues) and found nothing similar.

### Motivation

Apache Paimon Java has introduced the VECTOR data type based on [PIP-40](https://cwiki.apache.org/confluence/spaces/PAIMON/pages/399279132/PIP-40+Introduce+a+new+Vector+data+type). It supports schema representation, regular data-file storage such as Parquet, dedicated vector storage, reads, writes, and Data Evolution.

Paimon C++ currently recognizes .vector. file names for some file-level bookkeeping, but it does not yet provide a VECTOR logical type, Arrow mapping, serialization, storage, or end-to-end read and write support.

This issue implements the VECTOR roadmap item tracked in [#186](https://github.com/apache/paimon-cpp/issues/186).

### Solution

Introduce `VECTOR` support incrementally, while keeping the schema and storage behavior compatible with Apache Paimon Java.

#### Phase 1: Schema and regular Parquet storage

- [x] Add `VECTOR` to the Paimon C++ logical type system.
- [x] Implement schema JSON serialization and deserialization compatible with Paimon Java.
- [x] Map `VECTOR` to Arrow `FixedSizeList`.
- [x] Support the element types defined by PIP-40:
- `BOOLEAN`
- `TINYINT`
- `SMALLINT`
- `INT`
- `BIGINT`
- `FLOAT`
- `DOUBLE`
- [x] Validate that the dimension is positive and fixed.
- [x] Validate that the written vector length equals `N`.
- [x] Reject null vector elements.
- [x] Support reading and writing VECTOR columns in regular Parquet data files.
- [x] Add end-to-end append-table tests.
- [x] Add Java/C++ schema and file compatibility tests.

#### Phase 2: Schema Evolution, Data Evolution, and Primary-Key Table Baseline

##### Schema Evolution

- [x] Support adding and dropping VECTOR columns.
- [x] Support reading files written with previous table schemas.
- [x] Preserve VECTOR values when projecting files across schema versions.
- [x] Reject incompatible dimension changes, such as
`VECTOR` to `VECTOR`.
- [x] Reject incompatible element-type changes.
- [x] Add schema-evolution integration tests for both append-only and
supported primary-key tables.

##### Data Evolution Read/Write

- [ ] Support VECTOR columns in row-tracking append-only tables with
`data-evolution.enabled = true`.
- [ ] Support full-row writes containing VECTOR columns.
- [ ] Support partial-column writes containing VECTOR columns through the
write-schema path.
- [ ] Merge VECTOR values from files covering the same row-id range.
- [ ] Support reading VECTOR columns across different file schema IDs.
- [ ] Add Data Evolution integration tests covering full writes,
partial VECTOR writes, null values, and mixed old/new schema files.

##### Primary-Key Table Baseline

- [x] Allow VECTOR columns as ordinary non-key value columns in
deduplicate primary-key tables backed by regular Parquet files.
- [x] Support insert, same-key update, and read.
- [x] Support write-buffer spill and compaction.
- [x] Support adding and dropping VECTOR value columns while retaining
readability of files written with previous schemas.
- [x] Reject VECTOR columns as primary, partition, bucket, sequence,
sequence-group ordering, or sorting fields.
- [x] Explicitly reject lookup, partial-update, aggregation, and other
unsupported primary-key configurations until they are implemented.
- [x] Add end-to-end primary-key table tests for write, update, read,
spill, compaction, and schema evolution.
#### Phase 3: Dedicated vector storage

- [ ] Support the vector file format configuration used by Paimon Java.
- [ ] Support reading and writing dedicated `*.vector.vortex` files.
- [ ] Integrate vector files with row tracking and Data Evolution.
- [ ] Integrate vector files with scan planning, file commits, and conflict handling.
- [ ] Add Java, Python, and C++ Vortex compatibility tests.

### Initial scope

The initial implementation can focus on Phase 1, providing a usable end-to-end vertical slice through schema representation, Arrow mapping, and regular Parquet reads and writes.

Dedicated Vortex vector storage and Data Evolution can be delivered through follow-up pull requests under this issue.

The following items are not required for the initial implementation:

- ORC VECTOR support
- Vector indexes or similarity search
- Changing vector dimensions through schema evolution
- VECTOR values inside shared-shredding MAP columns
- Element types not supported by Apache Paimon Java

PIP-40 should only be considered fully supported after all phases are complete. Completing Phase 1 means that regular Parquet VECTOR storage is supported, but does not imply support for dedicated Vortex vector files.

### Anything else?

Related roadmap: [#186](https://github.com/apache/paimon-cpp/issues/186)

### Are you willing to submit a PR?

- [x] I'm willing to submit a PR!

コントリビューションガイド

コントリビューションガイドを開く

評価

この issue はまだ評価されていません。

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。