Table-model support for the Flink connectors: concrete design and three open questions
- Langage dominant
- Java
- Étoiles
- 6.4k
- Forks
- 1.2k
- Merge moyen
- 1 j 23 h
- PR mergées (30 j)
- 115
Description
## Concrete design: table-model support for the Flink connectors
Follow-up to [DISCUSS] Table-model support for the Flink connectors on
dev@iotdb.apache.org (2026-08-08). That thread received no replies, so what
follows is the design I proposed there written out concretely. The three
questions I asked on the list are still open, and I have kept them open here
rather than treating silence as agreement on any of them.
### The problem
`flink-sql-iotdb-connector`'s schema mapping is the tree model, not a
configuration of it. In `IoTDBSinkFunction` a Flink column name is parsed as an
IoTDB path and split into a device and a measurement:
```
:132-136 PathUtils.splitPathToDetachedNodes(fieldName);
measurement = nodes[nodes.length - 1];
device = join(copyOfRange(nodes, 0, nodes.length - 1), '.');
:108-113 session.insertAlignedRecord(...) / session.insertRecord(...)
:86 new Session.Builder().nodeUrls(..).username(..).password(..).build()
```
In table mode there is no path to split. A column is a TAG, FIELD or ATTRIBUTE
under `database.table`, and which of the three it is carries meaning a name
cannot express. The connector's option list agrees that this is not a
configuration gap: there is no `database` and no dialect option, and `aligned`
and `cdc.pattern` are tree concepts.
### Proposed shape
A new module `flink-iotdb-table-connector`, leaving `flink-sql-iotdb-connector`
untouched, mirroring how this repository already split Spark:
`spark-iotdb-connector` and `spark-iotdb-table-connector` are parallel trees with
their own parent poms, and the table one has its own `spark-iotdb-table-common`
rather than sharing the tree one's.
The objection to a separate module is duplicated CDC, lookup and bounded-scan
machinery. That objection applied equally to the Spark split and the project
accepted it there, so the cost is one this repository has already weighed for
this exact problem.
### Three questions that are still open
The DISCUSS thread drew no replies. Lazy consensus covers "nobody objected to
the direction"; it does not answer these, and one of them rests on reading I
explicitly flagged as incomplete.
1. **Is a separate `flink-iotdb-table-connector` the right shape here?** The
Spark precedent is the argument for it, but Spark's split may have had
reasons that do not carry over.
2. **Is the sink the right place to start?** My reading was that the source's
tree couplings are the dialect-less `Session` and the `TIME` clause — a
smaller and different problem from a mapping with no table-mode analogue —
which would put the design work in the sink. **I have not read the CDC or
lookup paths.** If the source has couplings I have not found, this ordering
is wrong and I would rather know before writing code than after.
3. **Is anyone already working on this?** I searched issues and pull requests in
both `iotdb-extras` and `iotdb` and found nothing on Flink and the table
model, but a search is not the same as asking.
### What I plan to do next
Subject to the above: build and run both existing Flink connectors first. My
DISCUSS post was explicit that everything in it came from reading source and
that I had run neither connector. That is a reasonable basis for proposing a
shape; it is not a reasonable basis for implementing one, so running them is the
first step rather than a later one.
I am happy to take the implementation if the direction holds, and equally happy
to hand the design to whoever is better placed to do it.
Guide de contribution
Ouvrir le guide de contribution
Piste de recherche
Commence par construire et exécuter les deux connecteurs Flink existants, puis lis IoTDBSinkFunction autour du parsing des chemins cité et des appels à Session. Compare les arborescences parallèles des connecteurs Spark et examine les chemins source, CDC et lookup avant de décider si un table connector séparé et une implémentation sink-first conviennent. C’est terminé lorsque les trois questions de conception ont une réponse et que la forme de l’implémentation est convenue.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Évaluation
- Stack technique
- java
- Domaine
- databases
- Type d'issue
- Fonctionnalité
- Difficulté
- 5/5
- Temps estimé
- Plus d'une semaine
- Activité
- Active
- Clarté
- À clarifier
- Accessibilité débutants
- 30/100