apache / apache/iggy

Server panics/aborts on EMFILE (Too many open files) during index rebuild instead of failing gracefully

Open
#3,645 3 comments 0 reactions 0 assignees View on GitHub
bug question server
Dominant language
Rust
Stars
4.9k
Forks
432
Avg merge
2d 10h
Merged PRs (30d)
173

Description

## Bug description

While running certain workloads against a locally started `iggy-server` (for example the Python SDK unit tests), the server can hit `Too many open files (os error 24)` and then **panic and abort the whole process** instead of degrading gracefully.

## Steps to reproduce

1. Start the server:

```bash
cargo run --bin iggy-server -- --with-default-root-credentials --fresh
```

2. Run the Python SDK tests (creating/dropping many streams and opening many connections in a short time):

```bash
uv run pytest tests/ -v
```

## Observed logs

```
2026-07-10T10:14:50.275123Z ERROR shard-0 server::tcp::tcp_listener: Unable to accept TCP socket. Too many open files (os error 24)
2026-07-10T10:14:50.275208Z INFO shard-0 server::shard::system::clients: Deleted tcp client with ID: 1128836228 for IP address: 127.0.0.1:57453
2026-07-10T10:14:50.275257Z INFO shard-0 server::tcp::tcp_listener: Successfully closed for client 1128836228, address 127.0.0.1:57453
2026-07-10T10:14:54.112973Z INFO shard-0 server::shard::tasks::periodic::sysinfo_printer: CPU: 14.89%/23.30% (IggyUsage/Total), Mem: 41.60%/254.35 MB/10.03 GB/17.18 GB (Free/IggyUsage/TotalUsed/Total), Disk: 12.53 GB/494.38 GB (Free/Total), IggyUsage: 37.88 KB, Clients: 18, Messages: 183, Read: 94.45 MB, Written: 7.18 MB
2026-07-10T10:14:55.276808Z INFO shard-0 server::tcp::tcp_listener: Accepted new TCP connection: 127.0.0.1:57461
2026-07-10T10:14:55.276912Z INFO shard-0 server::tcp::tcp_listener: Added tcp client with session: client ID: 1499973608, IP address: 127.0.0.1:57461 for IP address: 127.0.0.1:57461
2026-07-10T10:14:55.276961Z INFO shard-0 server::tcp::tcp_listener: Created new session: client ID: 1499973608, IP address: 127.0.0.1:57461
2026-07-10T10:14:55.277123Z ERROR shard-0 server::tcp::tcp_listener: Unable to accept TCP socket. Too many open files (os error 24)
2026-07-10T10:14:55.277593Z INFO shard-0 server::shard::system::clients: Deleted tcp client with ID: 1499973608 for IP address: 127.0.0.1:57461
2026-07-10T10:14:55.277774Z INFO shard-0 server::tcp::tcp_listener: Successfully closed for client 1499973608, address 127.0.0.1:57461
2026-07-10T10:15:00.278819Z INFO shard-0 server::tcp::tcp_listener: Accepted new TCP connection: 127.0.0.1:57483
2026-07-10T10:15:00.279041Z INFO shard-0 server::tcp::tcp_listener: Added tcp client with session: client ID: 3607919687, IP address: 127.0.0.1:57483 for IP address: 127.0.0.1:57483
2026-07-10T10:15:00.279106Z INFO shard-0 server::tcp::tcp_listener: Created new session: client ID: 3607919687, IP address: 127.0.0.1:57483
2026-07-10T10:15:00.279525Z ERROR shard-0 server::tcp::tcp_listener: Unable to accept TCP socket. Too many open files (os error 24)
2026-07-10T10:15:00.280162Z INFO shard-0 server::shard::system::clients: Deleted tcp client with ID: 3607919687 for IP address: 127.0.0.1:57483
2026-07-10T10:15:00.281297Z INFO shard-0 server::tcp::tcp_listener: Successfully closed for client 3607919687, address 127.0.0.1:57483
2026-07-10T10:15:00.289213Z INFO shard-0 server::streaming::streams::storage: Saved stream with ID: 232.
2026-07-10T10:15:00.291016Z INFO shard-0 server::streaming::topics::storage: Saved topic with ID: 0. for stream with ID: 232
2026-07-10T10:15:00.291271Z INFO shard-0 server::streaming::partitions::storage: Saving partition with ID: 0 for stream with ID: 232 and topic with ID: 0...
2026-07-10T10:15:00.291847Z INFO shard-0 server::streaming::partitions::storage: Saved partition with start ID: 0 for stream with ID: 232 and topic with ID: 0, path: local_data/streams/232/topics/0/partitions/0.
2026-07-10T10:15:00.292565Z INFO shard-3 server::shard::system::partitions: Initializing partition in local_partitions: partition ID: 0 for topic ID: 0 for stream ID: 232
2026-07-10T10:15:00.293123Z INFO shard-3 server::shard::system::partitions: No segments found on disk for partition ID: 0 for topic ID: 0 for stream ID: 232, creating initial segment
2026-07-10T10:15:00.293390Z ERROR shard-3 err_trail: Failed to open index file: local_data/streams/232/topics/0/partitions/0/00000000000000000000.index. Too many open files (os error 24)
2026-07-10T10:15:00.297908Z INFO shard-3 server::shard::system::partitions: Initializing partition in local_partitions: partition ID: 0 for topic ID: 0 for stream ID: 232
2026-07-10T10:15:00.298552Z WARN shard-3 server::bootstrap: Index at path local_data/streams/232/topics/0/partitions/0/00000000000000000000.index does not exist, rebuilding it based on local_data/streams/232/topics/0/partitions/0/00000000000000000000.log...

thread 'shard-3' (10092257) panicked at core/server/src/bootstrap.rs:226:17:
Failed to rebuild index for partition with ID: 0 for stream with ID: 232 and topic with ID: 0. Error: IO error
note: run with `RUST_BACKTRACE=1` environment variable to display a backtrace
[1] 76285 abort cargo run --bin iggy-server -- --with-default-root-credentials --fresh
```

## Guessed root cause

The trigger is file-descriptor exhaustion (`EMFILE` / `Too many open files`). The test suite opens/closes a lot of TCP connections and creates many streams/partitions rapidly, so the process runs into the (fairly low, on macOS by default) fd limit. So far this is expected pressure.

The problem is how that pressure is handled downstream. The chain seems to be:

1. `EMFILE` first shows up in the TCP accept loop (`server::tcp::tcp_listener`), where it is correctly handled: logged as an error and the accept loop keeps going (`tcp_listener.rs:143`).
2. But the same `EMFILE` also hits the storage path. When a new partition segment is created, opening the index file fails (`Failed to open index file: ... Too many open files`).
3. Because the index file could not be opened/created, the bootstrap logic concludes the index "does not exist" and tries to **rebuild** it from the log.
4. The rebuild itself also fails on I/O (still out of fds), and here the code panics instead of returning an error:

```rust
// core/server/src/bootstrap.rs:225
index_rebuilder.rebuild().await.unwrap_or_else(|e| {
panic!(
"Failed to rebuild index for partition with ID: {} for stream with ID: {} and topic with ID: {}. Error: {e}",
partition_id, stream_id, topic_id,
)
});
```

The panic on `shard-3` aborts the whole server process.

There also appear to be several nearby `.unwrap()` calls in the same function that assume infallible I/O (e.g. `try_exists(...).await.unwrap()` at `bootstrap.rs:208`), which would fail the same way under fd pressure. So a transient, recoverable OS condition (`EMFILE`) ends up taking the entire server down.

## Questions

- Is this the expected behavior? It seems undesirable that a transient `EMFILE` during normal client churn can panic and abort the whole server rather than failing just the offending operation.
- Would it make sense to propagate the error (return an `IggyError`) from `bootstrap` instead of `panic!` / `unwrap`, so the failing partition/stream operation fails gracefully while the server stays up?
- Should the storage layer distinguish "index genuinely missing" from "could not open index (I/O error)" so it does not attempt a rebuild in a situation where the real problem is fd exhaustion?
- Would it be worth raising/checking the fd `RLIMIT_NOFILE` at startup (and/or documenting the requirement), given the server is fd-hungry per shard/segment?

Happy to help with a PR if there is agreement on the preferred direction.

## Environment

- OS: macOS 26.5.1 (build 25F80), arm64 (Apple Silicon)
- iggy: `0.8.1-edge.1`, commit `8daf5464c94890e3a7fc87ac7851acdc5ed338b3` (server binary, `--with-default-root-credentials --fresh`)
- Rust toolchain: 1.96.0 (pinned in `rust-toolchain.toml`)
- Reproduced via Python SDK unit tests (`uv run pytest tests/ -v`)

Contributor guide

Open the contributing guide

Research direction

Start with core/server/src/bootstrap.rs around lines 208-226 and compare its I/O error handling with the accept loop in core/server/src/tcp/tcp_listener.rs:143. Reproduce with the documented cargo server command and uv run pytest tests/ -v under the stated file-descriptor pressure. Done means EMFILE does not panic or abort the server, and the affected operation reports an error while the process remains running.

Written by the indexing model from the issue text.

Assessment

Tech stack
rust
Domain
backend, databases
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Clearly specified
Newbie friendliness
48/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.