opendal FileIO hard-wires TimeoutLayer defaults (10s per-IO) with no way to configure them; large Parquet row groups fail deterministically on S3
- Dominant language
- Rust
- Stars
- 1.4k
- Forks
- 567
- Avg merge
- 2d 2h
- Merged PRs (30d)
- 93
Description
### Apache Iceberg Rust version
Reproduced on `iceberg` 0.10.1 / `iceberg-storage-opendal` 0.10.1. The code path is unchanged on `main` as of 2026-08-07 (`86d9d7d`).
### Describe the bug
`iceberg-storage-opendal` wraps every FileIO operator in `TimeoutLayer::new()`, and nothing in the crate's property or builder surface can override it:
https://github.com/apache/iceberg-rust/blob/86d9d7dc4297302496e6a121934de719c3ba9bea/crates/storage/opendal/src/lib.rs#L368
`TimeoutLayer::new()` uses opendal's defaults — **60 s overall, 10 s per IO operation** ([opendal 0.58 docs](https://docs.rs/opendal/0.58.0/opendal/layers/struct.TimeoutLayer.html)).
The 10 s `io_timeout` is the one that bites on the write path. opendal turns a single large `write` into a single S3 multipart part, and parquet-rs flushes a completed row group as one write — so in practice **the row group is the part**. Any row group whose encoded size exceeds what the writer's uplink can push in 10 s fails, and fails *deterministically*: the same request is retried by the `RetryLayer` above it and cannot ever fit inside the budget, so every attempt dies at the same place.
This is easy to hit on a table with a wide binary/string column, where row-group size is dominated by a handful of large values rather than by the row count the caller configured.
### To Reproduce
Environment: `iceberg` 0.10.1 + `iceberg-catalog-glue`, `parquet`/`arrow` 58, rustc 1.97.1, S3 over a WAN link (developer laptop, not in-region).
1. Create an Iceberg table with a `binary` payload column carrying large values — in our corpus ~9,200 rows / ~1.3 GB total, p50 ≈ 2.9 KB, p99 ≈ 1.4 MB, max ≈ 33 MB.
2. Write it through `ParquetWriterBuilder` with row-count row-group sizing, `WriterProperties::set_max_row_group_size(2048)`.
3. The write fails while flushing a row group:
```
Unexpected (persistent) at write, context: {
upload_id: , part_number: 0, service: s3,
path: /data/00000-....parquet,
size: 125382156, written: 428988086
} => write part timeout
1: External: External: Unexpected => Failure in doing io operation, ...
...
4: Unexpected (persistent) at write, context: { timeout: 10 } => io operation timeout reached
```
`size: 125382156` is the part being flushed; `timeout: 10` is the hard-wired `io_timeout`. **Three separate runs of the identical configuration failed at the identical byte offset.**
Two controls isolate the cause:
* **Same data, same row-group configuration, written with `arrow-rs` + `object_store` directly** (no Iceberg FileIO): completed in 44 s, no error. So this is not parquet, not arrow, and not S3.
* **Same data through Iceberg with smaller row groups** (512 rows, ≈4× smaller parts; and byte-based row-group targets of 4/16/64 MiB set in code): all succeed.
The threshold itself is bandwidth-dependent — an in-region task will tolerate a much larger row group than this laptop did — but the *ceiling exists at every bandwidth* and is not visible or adjustable from the Iceberg API.
### Expected behavior
Either of:
1. **Make the timeouts configurable.** A FileIO property in the existing `client.*` namespace (alongside `client.region`, `client.assume-role.arn`, …) — e.g. `client.io-timeout-ms` / `client.timeout-ms` — or a builder hook on the storage config, so a caller writing large parts can raise the per-IO budget.
2. **Scale the per-IO timeout with the size of the write being attempted**, so a large multipart part gets a proportional budget instead of a fixed 10 s.
The current default is a sensible one, and the comment at the call site is right that `TimeoutLayer` must sit inside `RetryLayer` so each attempt is independently bounded. The problem is only that it is a *hard* ceiling: it silently caps the maximum row-group size a table can be written with, as a function of the writer's uplink bandwidth rather than of anything the table or the data expresses — and it surfaces as a persistent, retry-proof write failure rather than as a configuration error.
### Willingness to contribute
I cannot contribute a fix for this bug at this time.
Contributor guide
Research direction
Start at crates/storage/opendal/src/lib.rs around line 368 and trace how TimeoutLayer is constructed relative to the existing client properties and retry layer. Run the supplied large-row-group S3 reproduction, then verify that the chosen configuration or sizing behavior allows the affected write to complete without the fixed 10-second per-IO ceiling.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- aws, rust
- Domain
- backend, cloud
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Mostly clear
- Newbie friendliness
- 48/100