Table-model support for the Flink connectors: concrete design and three open questions
- Lenguaje dominante
- Java
- Estrellas
- 6.4k
- Forks
- 1.2k
- Merge medio
- 1 d 23 h
- PR fusionados (30 d)
- 115
Descripción
## Concrete design: table-model support for the Flink connectors
Follow-up to [DISCUSS] Table-model support for the Flink connectors on
dev@iotdb.apache.org (2026-08-08). That thread received no replies, so what
follows is the design I proposed there written out concretely. The three
questions I asked on the list are still open, and I have kept them open here
rather than treating silence as agreement on any of them.
### The problem
`flink-sql-iotdb-connector`'s schema mapping is the tree model, not a
configuration of it. In `IoTDBSinkFunction` a Flink column name is parsed as an
IoTDB path and split into a device and a measurement:
```
:132-136 PathUtils.splitPathToDetachedNodes(fieldName);
measurement = nodes[nodes.length - 1];
device = join(copyOfRange(nodes, 0, nodes.length - 1), '.');
:108-113 session.insertAlignedRecord(...) / session.insertRecord(...)
:86 new Session.Builder().nodeUrls(..).username(..).password(..).build()
```
In table mode there is no path to split. A column is a TAG, FIELD or ATTRIBUTE
under `database.table`, and which of the three it is carries meaning a name
cannot express. The connector's option list agrees that this is not a
configuration gap: there is no `database` and no dialect option, and `aligned`
and `cdc.pattern` are tree concepts.
### Proposed shape
A new module `flink-iotdb-table-connector`, leaving `flink-sql-iotdb-connector`
untouched, mirroring how this repository already split Spark:
`spark-iotdb-connector` and `spark-iotdb-table-connector` are parallel trees with
their own parent poms, and the table one has its own `spark-iotdb-table-common`
rather than sharing the tree one's.
The objection to a separate module is duplicated CDC, lookup and bounded-scan
machinery. That objection applied equally to the Spark split and the project
accepted it there, so the cost is one this repository has already weighed for
this exact problem.
### Three questions that are still open
The DISCUSS thread drew no replies. Lazy consensus covers "nobody objected to
the direction"; it does not answer these, and one of them rests on reading I
explicitly flagged as incomplete.
1. **Is a separate `flink-iotdb-table-connector` the right shape here?** The
Spark precedent is the argument for it, but Spark's split may have had
reasons that do not carry over.
2. **Is the sink the right place to start?** My reading was that the source's
tree couplings are the dialect-less `Session` and the `TIME` clause — a
smaller and different problem from a mapping with no table-mode analogue —
which would put the design work in the sink. **I have not read the CDC or
lookup paths.** If the source has couplings I have not found, this ordering
is wrong and I would rather know before writing code than after.
3. **Is anyone already working on this?** I searched issues and pull requests in
both `iotdb-extras` and `iotdb` and found nothing on Flink and the table
model, but a search is not the same as asking.
### What I plan to do next
Subject to the above: build and run both existing Flink connectors first. My
DISCUSS post was explicit that everything in it came from reading source and
that I had run neither connector. That is a reasonable basis for proposing a
shape; it is not a reasonable basis for implementing one, so running them is the
first step rather than a later one.
I am happy to take the implementation if the direction holds, and equally happy
to hand the design to whoever is better placed to do it.
Guía de contribución
Línea de trabajo
Empieza compilando y ejecutando ambos conectores de Flink existentes, y después lee IoTDBSinkFunction en torno al análisis de rutas citado y las llamadas a Session. Compara los árboles paralelos de conectores de Spark e inspecciona las rutas de source, CDC y lookup antes de decidir si encajan un table connector separado y una implementación sink-first. Se considera terminado cuando las tres preguntas de diseño tienen respuesta y se ha acordado la forma de la implementación.
Escrito por el modelo de indexación a partir del texto del issue.
Evaluación
- Stack tecnológico
- java
- Área
- databases
- Tipo de issue
- Nueva funcionalidad
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Estado de actividad
- Activo
- Claridad
- Necesita aclaración
- Aptitud para principiantes
- 30/100