apache / apache/iotdb

Table-model support for the Flink connectors: concrete design and three open questions

オープン
#18,565 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る
主要言語
Java
スター
6.4k
フォーク
1.2k
平均マージ
1日 23時間
マージ済み PR(30日)
115

説明

## Concrete design: table-model support for the Flink connectors

Follow-up to [DISCUSS] Table-model support for the Flink connectors on
dev@iotdb.apache.org (2026-08-08). That thread received no replies, so what
follows is the design I proposed there written out concretely. The three
questions I asked on the list are still open, and I have kept them open here
rather than treating silence as agreement on any of them.

### The problem

`flink-sql-iotdb-connector`'s schema mapping is the tree model, not a
configuration of it. In `IoTDBSinkFunction` a Flink column name is parsed as an
IoTDB path and split into a device and a measurement:

```
:132-136 PathUtils.splitPathToDetachedNodes(fieldName);
measurement = nodes[nodes.length - 1];
device = join(copyOfRange(nodes, 0, nodes.length - 1), '.');
:108-113 session.insertAlignedRecord(...) / session.insertRecord(...)
:86 new Session.Builder().nodeUrls(..).username(..).password(..).build()
```

In table mode there is no path to split. A column is a TAG, FIELD or ATTRIBUTE
under `database.table`, and which of the three it is carries meaning a name
cannot express. The connector's option list agrees that this is not a
configuration gap: there is no `database` and no dialect option, and `aligned`
and `cdc.pattern` are tree concepts.

### Proposed shape

A new module `flink-iotdb-table-connector`, leaving `flink-sql-iotdb-connector`
untouched, mirroring how this repository already split Spark:
`spark-iotdb-connector` and `spark-iotdb-table-connector` are parallel trees with
their own parent poms, and the table one has its own `spark-iotdb-table-common`
rather than sharing the tree one's.

The objection to a separate module is duplicated CDC, lookup and bounded-scan
machinery. That objection applied equally to the Spark split and the project
accepted it there, so the cost is one this repository has already weighed for
this exact problem.

### Three questions that are still open

The DISCUSS thread drew no replies. Lazy consensus covers "nobody objected to
the direction"; it does not answer these, and one of them rests on reading I
explicitly flagged as incomplete.

1. **Is a separate `flink-iotdb-table-connector` the right shape here?** The
Spark precedent is the argument for it, but Spark's split may have had
reasons that do not carry over.

2. **Is the sink the right place to start?** My reading was that the source's
tree couplings are the dialect-less `Session` and the `TIME` clause — a
smaller and different problem from a mapping with no table-mode analogue —
which would put the design work in the sink. **I have not read the CDC or
lookup paths.** If the source has couplings I have not found, this ordering
is wrong and I would rather know before writing code than after.

3. **Is anyone already working on this?** I searched issues and pull requests in
both `iotdb-extras` and `iotdb` and found nothing on Flink and the table
model, but a search is not the same as asking.

### What I plan to do next

Subject to the above: build and run both existing Flink connectors first. My
DISCUSS post was explicit that everything in it came from reading source and
that I had run neither connector. That is a reasonable basis for proposing a
shape; it is not a reasonable basis for implementing one, so running them is the
first step rather than a later one.

I am happy to take the implementation if the direction holds, and equally happy
to hand the design to whoever is better placed to do it.

コントリビューションガイド

コントリビューションガイドを開く

調査の方向性

まず既存の2つの Flink コネクタをビルドして実行し、続いて引用されているパス解析と Session 呼び出しの周辺にある IoTDBSinkFunction を読みます。対応する Spark コネクタのツリーを比較し、source、CDC、lookup の各パスを調べてから、個別の table connector と sink-first 実装が適合するかどうかを判断します。3つの設計上の問いに答えが出て、実装の形について合意できれば完了です。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
java
領域
databases
issue の種類
機能追加
難易度
5/5
見積もり時間
1週間以上
活発さ
活発
明瞭さ
説明が足りない
初心者へのやさしさ
30/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。