[Bug] File descriptor leak in CLI CSV import: CSVParser from readCsvFile is never closed
- 主要言語
- Java
- スター
- 6.4k
- フォーク
- 1.2k
- 平均マージ
- 1日 23時間
- マージ済み PR(30日)
- 115
説明
### Search before asking
- [x] I searched in the [issues](https://github.com/apache/iotdb/issues) and found nothing similar.
### Version
`master` (2.0.x). The affected code is also present in released 2.0.x versions.
### Describe the bug and provide the minimal reproduce step
`AbstractDataTool.readCsvFile(String)` builds a `CSVParser` over `new InputStreamReader(new FileInputStream(path))` and returns it. The `CSVParser` owns that `FileInputStream`, but the CLI import code paths that call it never close the returned parser — it is assigned to a local inside a plain `try { ... }` block (no try-with-resources, no `finally`). Every imported file therefore leaks its file descriptor, and the early `return`s for an empty file or an invalid header leak it immediately, because the parser is opened before those checks run.
Affected call sites (current `master`, also present in released 2.0.x):
- `iotdb-client/cli/src/main/java/org/apache/iotdb/tool/data/ImportData.java` — `importFromSingleFile`
- `iotdb-client/cli/src/main/java/org/apache/iotdb/tool/data/ImportDataTree.java` — `importFromCsvFile`
- `iotdb-client/cli/src/main/java/org/apache/iotdb/tool/data/ImportDataTable.java` — `importFromCsvFile`
- `iotdb-client/cli/src/main/java/org/apache/iotdb/tool/schema/ImportSchemaTree.java` — `importSchemaFromCsvFile` (this class has its own copy of `readCsvFile`)
Minimal reproduce step:
1. Create a directory containing a large number of small CSV files — more than the process open-file limit (for example a few thousand files under `ulimit -n 1024`).
2. Run the CLI data import over that directory.
3. The import fails partway through with `Too many open files`. (A single import already leaks one descriptor; it is simply not fatal until enough accumulate.)
### What did you expect to see?
Each `CSVParser` (and the `FileInputStream` it wraps) is closed after the file is processed, on every exit path — including the empty-file / invalid-header early returns. Importing a large directory of CSV files should not exhaust the process's file descriptors.
### What did you see instead?
The `CSVParser` returned by `readCsvFile` is never closed, so its underlying `FileInputStream` stays open. Importing a directory with enough CSV files leaks descriptors until the import fails with `Too many open files`.
### Anything else?
The record `Stream` is fully consumed inside the same block before the method returns, so consuming each parser in a try-with-resources (closing it on scope exit) is safe.
### Are you willing to submit a PR?
- [x] I'm willing to submit a PR!
コントリビューションガイド
調査の方向性
まず AbstractDataTool.readCsvFile と、指定されたインポートエントリポイントである ImportData.importFromSingleFile、ImportDataTree.importFromCsvFile、ImportDataTable.importFromCsvFile、ImportSchemaTree.importSchemaFromCsvFile を読みます。低い ulimit の下で多数の小さな CSV ファイルを使って問題を再現し、その後、正常終了時と早期 return の場合のどちらでも、ファイルディスクリプタを使い果たすことなくすべてのパーサーがクローズされることを確認します。
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- java
- 領域
- cli
- issue の種類
- バグ
- 難易度
- 3/5
- 見積もり時間
- 1〜2日
- 活発さ
- 静か
- 明瞭さ
- 明確に書かれている
- 初心者へのやさしさ
- 76/100