apache / apache/iotdb

[Bug] File descriptor leak in CLI CSV import: CSVParser from readCsvFile is never closed

未关闭
#18,237 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
主要语言
Java
星标
6.4k
派生
1.2k
平均合并
1 天 23 小时
30 天内合并 PR
115

描述

### Search before asking

- [x] I searched in the [issues](https://github.com/apache/iotdb/issues) and found nothing similar.

### Version

`master` (2.0.x). The affected code is also present in released 2.0.x versions.

### Describe the bug and provide the minimal reproduce step

`AbstractDataTool.readCsvFile(String)` builds a `CSVParser` over `new InputStreamReader(new FileInputStream(path))` and returns it. The `CSVParser` owns that `FileInputStream`, but the CLI import code paths that call it never close the returned parser — it is assigned to a local inside a plain `try { ... }` block (no try-with-resources, no `finally`). Every imported file therefore leaks its file descriptor, and the early `return`s for an empty file or an invalid header leak it immediately, because the parser is opened before those checks run.

Affected call sites (current `master`, also present in released 2.0.x):

- `iotdb-client/cli/src/main/java/org/apache/iotdb/tool/data/ImportData.java` — `importFromSingleFile`
- `iotdb-client/cli/src/main/java/org/apache/iotdb/tool/data/ImportDataTree.java` — `importFromCsvFile`
- `iotdb-client/cli/src/main/java/org/apache/iotdb/tool/data/ImportDataTable.java` — `importFromCsvFile`
- `iotdb-client/cli/src/main/java/org/apache/iotdb/tool/schema/ImportSchemaTree.java` — `importSchemaFromCsvFile` (this class has its own copy of `readCsvFile`)

Minimal reproduce step:

1. Create a directory containing a large number of small CSV files — more than the process open-file limit (for example a few thousand files under `ulimit -n 1024`).
2. Run the CLI data import over that directory.
3. The import fails partway through with `Too many open files`. (A single import already leaks one descriptor; it is simply not fatal until enough accumulate.)

### What did you expect to see?

Each `CSVParser` (and the `FileInputStream` it wraps) is closed after the file is processed, on every exit path — including the empty-file / invalid-header early returns. Importing a large directory of CSV files should not exhaust the process's file descriptors.

### What did you see instead?

The `CSVParser` returned by `readCsvFile` is never closed, so its underlying `FileInputStream` stays open. Importing a directory with enough CSV files leaks descriptors until the import fails with `Too many open files`.

### Anything else?

The record `Stream` is fully consumed inside the same block before the method returns, so consuming each parser in a try-with-resources (closing it on scope exit) is safe.

### Are you willing to submit a PR?

- [x] I'm willing to submit a PR!

贡献指南

打开贡献指南

调研方向

首先阅读 AbstractDataTool.readCsvFile 以及指定的导入入口:ImportData.importFromSingleFile、ImportDataTree.importFromCsvFile、ImportDataTable.importFromCsvFile 和 ImportSchemaTree.importSchemaFromCsvFile。在较低的 ulimit 下使用许多小型 CSV 文件重现该问题,然后验证每个解析器在正常完成和提前 return 时都会被关闭,且不会耗尽文件描述符。

由索引模型根据 Issue 内容生成。

评估

技术栈
java
领域
cli
Issue 类型
缺陷
难度
3/5
预计耗时
1-2 天
活跃度
冷清
描述清晰度
描述清楚
新手友好度
76/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。