aws / aws/aws-sdk-ruby

Object#upload_stream retains unbounded memory when the source outpaces the upload

Open
#3,393 1 comment 0 reactions 0 assignees View on GitHub
bug
Dominant language
Ruby
Stars
3.7k
Forks
1.2k
Avg merge
1d 4h
Merged PRs (30d)
5

Description

### Version

aws-sdk-s3 1.225.0 (Ruby 3.4)

### Problem

`Object#upload_stream` holds memory proportional to the *whole* object when the block writes faster than parts upload: a fast local source feeding a slower S3 sink. That defeats the purpose of a streaming upload.

### Cause

`MultipartStreamUploader#upload_with_executor` reads each part into a `StringIO` (`read_to_part_body`) and submits it through `DefaultExecutor#post`, which appends to an unbounded `Queue` and returns immediately (non-blocking). The internal `IO.pipe` backpressures the *writer* against the *reader*, but the reader then offloads into that unbounded queue, so nothing bounds read-ahead to the `~max_threads` parts that are actually uploading. Unsent 5 MB `StringIO`s accumulate up to roughly object size minus bytes already uploaded.

`tempfile: true` masks it (parts spill to disk), but the default is in-memory and unbounded.

### Impact

Streaming a multi-GB on-disk file spikes RSS by GBs. Concretely, Active Storage's `S3Service` archives large files via `upload_stream { |out| IO.copy_stream(file, out) }`; a 12 GB file drove ~3.6 GB of RSS and OOM'd a 4 GB worker. `upload_file` (lazy `FilePart`, reads from disk on demand) is unaffected under the same conditions; that divergence is the tell.

### Suggested fix

Bound read-ahead to the worker count: a `SizedQueue`, or have `post` block while all workers are busy. (`DefaultExecutor` arrived with the #3302 executor refactor; #1824 previously tuned `upload_stream` memory.)

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.