apache / apache/iotdb

Table-model support for the Flink connectors: concrete design and three open questions

Aperta
#18,565 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub
Lingua principale
Java
Stelle
6.4k
Fork
1.2k
Merge medio
1g 23h
PR unite (30g)
115

Descrizione

## Concrete design: table-model support for the Flink connectors

Follow-up to [DISCUSS] Table-model support for the Flink connectors on
dev@iotdb.apache.org (2026-08-08). That thread received no replies, so what
follows is the design I proposed there written out concretely. The three
questions I asked on the list are still open, and I have kept them open here
rather than treating silence as agreement on any of them.

### The problem

`flink-sql-iotdb-connector`'s schema mapping is the tree model, not a
configuration of it. In `IoTDBSinkFunction` a Flink column name is parsed as an
IoTDB path and split into a device and a measurement:

```
:132-136 PathUtils.splitPathToDetachedNodes(fieldName);
measurement = nodes[nodes.length - 1];
device = join(copyOfRange(nodes, 0, nodes.length - 1), '.');
:108-113 session.insertAlignedRecord(...) / session.insertRecord(...)
:86 new Session.Builder().nodeUrls(..).username(..).password(..).build()
```

In table mode there is no path to split. A column is a TAG, FIELD or ATTRIBUTE
under `database.table`, and which of the three it is carries meaning a name
cannot express. The connector's option list agrees that this is not a
configuration gap: there is no `database` and no dialect option, and `aligned`
and `cdc.pattern` are tree concepts.

### Proposed shape

A new module `flink-iotdb-table-connector`, leaving `flink-sql-iotdb-connector`
untouched, mirroring how this repository already split Spark:
`spark-iotdb-connector` and `spark-iotdb-table-connector` are parallel trees with
their own parent poms, and the table one has its own `spark-iotdb-table-common`
rather than sharing the tree one's.

The objection to a separate module is duplicated CDC, lookup and bounded-scan
machinery. That objection applied equally to the Spark split and the project
accepted it there, so the cost is one this repository has already weighed for
this exact problem.

### Three questions that are still open

The DISCUSS thread drew no replies. Lazy consensus covers "nobody objected to
the direction"; it does not answer these, and one of them rests on reading I
explicitly flagged as incomplete.

1. **Is a separate `flink-iotdb-table-connector` the right shape here?** The
Spark precedent is the argument for it, but Spark's split may have had
reasons that do not carry over.

2. **Is the sink the right place to start?** My reading was that the source's
tree couplings are the dialect-less `Session` and the `TIME` clause — a
smaller and different problem from a mapping with no table-mode analogue —
which would put the design work in the sink. **I have not read the CDC or
lookup paths.** If the source has couplings I have not found, this ordering
is wrong and I would rather know before writing code than after.

3. **Is anyone already working on this?** I searched issues and pull requests in
both `iotdb-extras` and `iotdb` and found nothing on Flink and the table
model, but a search is not the same as asking.

### What I plan to do next

Subject to the above: build and run both existing Flink connectors first. My
DISCUSS post was explicit that everything in it came from reading source and
that I had run neither connector. That is a reasonable basis for proposing a
shape; it is not a reasonable basis for implementing one, so running them is the
first step rather than a later one.

I am happy to take the implementation if the direction holds, and equally happy
to hand the design to whoever is better placed to do it.

Guida per i contributori

Apri la guida per i contributori

Direzione di ricerca

Inizia compilando ed eseguendo entrambi i connettori Flink esistenti, quindi leggi IoTDBSinkFunction intorno al parsing dei percorsi citato e alle chiamate a Session. Confronta gli alberi paralleli dei connettori Spark ed esamina i percorsi source, CDC e lookup prima di decidere se siano adatti un table connector separato e un’implementazione sink-first. Il lavoro è completato quando le tre domande di progettazione hanno una risposta e la forma dell’implementazione è stata concordata.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
java
Ambito
databases
Tipo di issue
Funzionalità
Difficoltà
5/5
Tempo stimato
Più di una settimana
Stato di attività
Attiva
Chiarezza
Da chiarire
Idoneità per principianti
30/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.