apache / apache/iotdb

[Bug] File descriptor leak in CLI CSV import: CSVParser from readCsvFile is never closed

Ouverte
#18,237 0 commentaires 0 réactions 0 personnes assignées Voir sur GitHub
Langage dominant
Java
Étoiles
6.4k
Forks
1.2k
Merge moyen
1 j 23 h
PR mergées (30 j)
115

Description

### Search before asking

- [x] I searched in the [issues](https://github.com/apache/iotdb/issues) and found nothing similar.

### Version

`master` (2.0.x). The affected code is also present in released 2.0.x versions.

### Describe the bug and provide the minimal reproduce step

`AbstractDataTool.readCsvFile(String)` builds a `CSVParser` over `new InputStreamReader(new FileInputStream(path))` and returns it. The `CSVParser` owns that `FileInputStream`, but the CLI import code paths that call it never close the returned parser — it is assigned to a local inside a plain `try { ... }` block (no try-with-resources, no `finally`). Every imported file therefore leaks its file descriptor, and the early `return`s for an empty file or an invalid header leak it immediately, because the parser is opened before those checks run.

Affected call sites (current `master`, also present in released 2.0.x):

- `iotdb-client/cli/src/main/java/org/apache/iotdb/tool/data/ImportData.java` — `importFromSingleFile`
- `iotdb-client/cli/src/main/java/org/apache/iotdb/tool/data/ImportDataTree.java` — `importFromCsvFile`
- `iotdb-client/cli/src/main/java/org/apache/iotdb/tool/data/ImportDataTable.java` — `importFromCsvFile`
- `iotdb-client/cli/src/main/java/org/apache/iotdb/tool/schema/ImportSchemaTree.java` — `importSchemaFromCsvFile` (this class has its own copy of `readCsvFile`)

Minimal reproduce step:

1. Create a directory containing a large number of small CSV files — more than the process open-file limit (for example a few thousand files under `ulimit -n 1024`).
2. Run the CLI data import over that directory.
3. The import fails partway through with `Too many open files`. (A single import already leaks one descriptor; it is simply not fatal until enough accumulate.)

### What did you expect to see?

Each `CSVParser` (and the `FileInputStream` it wraps) is closed after the file is processed, on every exit path — including the empty-file / invalid-header early returns. Importing a large directory of CSV files should not exhaust the process's file descriptors.

### What did you see instead?

The `CSVParser` returned by `readCsvFile` is never closed, so its underlying `FileInputStream` stays open. Importing a directory with enough CSV files leaks descriptors until the import fails with `Too many open files`.

### Anything else?

The record `Stream` is fully consumed inside the same block before the method returns, so consuming each parser in a try-with-resources (closing it on scope exit) is safe.

### Are you willing to submit a PR?

- [x] I'm willing to submit a PR!

Guide de contribution

Ouvrir le guide de contribution

Piste de recherche

Commencez par lire AbstractDataTool.readCsvFile et les points d’entrée d’importation indiqués : ImportData.importFromSingleFile, ImportDataTree.importFromCsvFile, ImportDataTable.importFromCsvFile et ImportSchemaTree.importSchemaFromCsvFile. Reproduisez le problème avec de nombreux petits fichiers CSV sous un ulimit faible, puis vérifiez que chaque analyseur est fermé en cas de fin normale et de retours anticipés, sans épuiser les descripteurs de fichiers.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
java
Domaine
cli
Type d'issue
Bug
Difficulté
3/5
Temps estimé
1-2 jours
Activité
Calme
Clarté
Clairement spécifiée
Accessibilité débutants
76/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.