Table-model support for the Flink connectors: concrete design and three open questions
- Lingua principale
- Java
- Stelle
- 6.4k
- Fork
- 1.2k
- Merge medio
- 1g 23h
- PR unite (30g)
- 115
Descrizione
## Concrete design: table-model support for the Flink connectors
Follow-up to [DISCUSS] Table-model support for the Flink connectors on
dev@iotdb.apache.org (2026-08-08). That thread received no replies, so what
follows is the design I proposed there written out concretely. The three
questions I asked on the list are still open, and I have kept them open here
rather than treating silence as agreement on any of them.
### The problem
`flink-sql-iotdb-connector`'s schema mapping is the tree model, not a
configuration of it. In `IoTDBSinkFunction` a Flink column name is parsed as an
IoTDB path and split into a device and a measurement:
```
:132-136 PathUtils.splitPathToDetachedNodes(fieldName);
measurement = nodes[nodes.length - 1];
device = join(copyOfRange(nodes, 0, nodes.length - 1), '.');
:108-113 session.insertAlignedRecord(...) / session.insertRecord(...)
:86 new Session.Builder().nodeUrls(..).username(..).password(..).build()
```
In table mode there is no path to split. A column is a TAG, FIELD or ATTRIBUTE
under `database.table`, and which of the three it is carries meaning a name
cannot express. The connector's option list agrees that this is not a
configuration gap: there is no `database` and no dialect option, and `aligned`
and `cdc.pattern` are tree concepts.
### Proposed shape
A new module `flink-iotdb-table-connector`, leaving `flink-sql-iotdb-connector`
untouched, mirroring how this repository already split Spark:
`spark-iotdb-connector` and `spark-iotdb-table-connector` are parallel trees with
their own parent poms, and the table one has its own `spark-iotdb-table-common`
rather than sharing the tree one's.
The objection to a separate module is duplicated CDC, lookup and bounded-scan
machinery. That objection applied equally to the Spark split and the project
accepted it there, so the cost is one this repository has already weighed for
this exact problem.
### Three questions that are still open
The DISCUSS thread drew no replies. Lazy consensus covers "nobody objected to
the direction"; it does not answer these, and one of them rests on reading I
explicitly flagged as incomplete.
1. **Is a separate `flink-iotdb-table-connector` the right shape here?** The
Spark precedent is the argument for it, but Spark's split may have had
reasons that do not carry over.
2. **Is the sink the right place to start?** My reading was that the source's
tree couplings are the dialect-less `Session` and the `TIME` clause — a
smaller and different problem from a mapping with no table-mode analogue —
which would put the design work in the sink. **I have not read the CDC or
lookup paths.** If the source has couplings I have not found, this ordering
is wrong and I would rather know before writing code than after.
3. **Is anyone already working on this?** I searched issues and pull requests in
both `iotdb-extras` and `iotdb` and found nothing on Flink and the table
model, but a search is not the same as asking.
### What I plan to do next
Subject to the above: build and run both existing Flink connectors first. My
DISCUSS post was explicit that everything in it came from reading source and
that I had run neither connector. That is a reasonable basis for proposing a
shape; it is not a reasonable basis for implementing one, so running them is the
first step rather than a later one.
I am happy to take the implementation if the direction holds, and equally happy
to hand the design to whoever is better placed to do it.
Guida per i contributori
Apri la guida per i contributori
Direzione di ricerca
Inizia compilando ed eseguendo entrambi i connettori Flink esistenti, quindi leggi IoTDBSinkFunction intorno al parsing dei percorsi citato e alle chiamate a Session. Confronta gli alberi paralleli dei connettori Spark ed esamina i percorsi source, CDC e lookup prima di decidere se siano adatti un table connector separato e un’implementazione sink-first. Il lavoro è completato quando le tre domande di progettazione hanno una risposta e la forma dell’implementazione è stata concordata.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Stack tecnologico
- java
- Ambito
- databases
- Tipo di issue
- Funzionalità
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Stato di attività
- Attiva
- Chiarezza
- Da chiarire
- Idoneità per principianti
- 30/100