[Bug] File descriptor leak in CLI CSV import: CSVParser from readCsvFile is never closed
- Ngôn ngữ chính
- Java
- Star
- 6.4k
- Fork
- 1.2k
- Merge trung bình
- 1 ngày 23 giờ
- Pull request đã merge (30 ngày)
- 115
Mô tả
### Search before asking
- [x] I searched in the [issues](https://github.com/apache/iotdb/issues) and found nothing similar.
### Version
`master` (2.0.x). The affected code is also present in released 2.0.x versions.
### Describe the bug and provide the minimal reproduce step
`AbstractDataTool.readCsvFile(String)` builds a `CSVParser` over `new InputStreamReader(new FileInputStream(path))` and returns it. The `CSVParser` owns that `FileInputStream`, but the CLI import code paths that call it never close the returned parser — it is assigned to a local inside a plain `try { ... }` block (no try-with-resources, no `finally`). Every imported file therefore leaks its file descriptor, and the early `return`s for an empty file or an invalid header leak it immediately, because the parser is opened before those checks run.
Affected call sites (current `master`, also present in released 2.0.x):
- `iotdb-client/cli/src/main/java/org/apache/iotdb/tool/data/ImportData.java` — `importFromSingleFile`
- `iotdb-client/cli/src/main/java/org/apache/iotdb/tool/data/ImportDataTree.java` — `importFromCsvFile`
- `iotdb-client/cli/src/main/java/org/apache/iotdb/tool/data/ImportDataTable.java` — `importFromCsvFile`
- `iotdb-client/cli/src/main/java/org/apache/iotdb/tool/schema/ImportSchemaTree.java` — `importSchemaFromCsvFile` (this class has its own copy of `readCsvFile`)
Minimal reproduce step:
1. Create a directory containing a large number of small CSV files — more than the process open-file limit (for example a few thousand files under `ulimit -n 1024`).
2. Run the CLI data import over that directory.
3. The import fails partway through with `Too many open files`. (A single import already leaks one descriptor; it is simply not fatal until enough accumulate.)
### What did you expect to see?
Each `CSVParser` (and the `FileInputStream` it wraps) is closed after the file is processed, on every exit path — including the empty-file / invalid-header early returns. Importing a large directory of CSV files should not exhaust the process's file descriptors.
### What did you see instead?
The `CSVParser` returned by `readCsvFile` is never closed, so its underlying `FileInputStream` stays open. Importing a directory with enough CSV files leaks descriptors until the import fails with `Too many open files`.
### Anything else?
The record `Stream` is fully consumed inside the same block before the method returns, so consuming each parser in a try-with-resources (closing it on scope exit) is safe.
### Are you willing to submit a PR?
- [x] I'm willing to submit a PR!
Hướng dẫn đóng góp
Hướng nghiên cứu
Bắt đầu bằng việc đọc AbstractDataTool.readCsvFile và các điểm vào import được nêu tên: ImportData.importFromSingleFile, ImportDataTree.importFromCsvFile, ImportDataTable.importFromCsvFile và ImportSchemaTree.importSchemaFromCsvFile. Tái hiện vấn đề với nhiều tệp CSV nhỏ dưới một ulimit thấp, sau đó xác minh rằng mọi parser đều được đóng khi hoàn tất bình thường và khi return sớm, mà không làm cạn kiệt các bộ mô tả tệp.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Đánh giá
- Công nghệ
- java
- Lĩnh vực
- cli
- Loại issue
- Lỗi
- Độ khó
- 3/5
- Thời gian dự kiến
- 1-2 ngày
- Mức độ hoạt động
- Ít trao đổi
- Độ rõ ràng
- Đặc tả rõ ràng
- Mức phù hợp với người mới
- 76/100