nextflow-io / nextflow-io/nextflow

S3ObjectSummaryLookup lists without a delimiter, over-listing sibling prefixes

Open Beginner friendly
#7,224 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

storage/aws
Dominant language
Groovy
Stars
3.5k
Forks
811
Avg merge
2d 11h
Merged PRs (30d)
61

Description

Bug report

Expected behavior and actual behavior

When resolving an S3 directory path (Files.exists / isDirectory / nf-schema format: path), S3ObjectSummaryLookup.lookup() lists by the bare key with no delimiter. Since S3Path strips the trailing slash, s3://bucket/sampleA/ is listed as prefix="sampleA", which also matches siblings like sampleA-extra/. matchName() rejects them afterward, but only after every sibling key is paginated at maxKeys=250.

On large layouts this inflates head-node S3 concurrency enough to reliably trigger the upstream CRT progress-accounting crash (aws/aws-sdk-java-v2#4790), which kills the event-loop thread and leaves the head job hung in submitted state with an empty .nextflow.log.

This is the only nf-amazon listing path missing the delimiter; S3Iterator.buildRequest() uses .prefix(key).delimiter("/").

// S3ObjectSummaryLookup.lookup()
request.prefix(s3Path.getKey());   // bare key, slash stripped
request.maxKeys(250);
// no .delimiter("/")
Steps to reproduce the problem
  1. Create sibling prefixes where one is a leading substring of the other, e.g. s3://mybucket/sampleA/ and s3://mybucket/sampleA-extra/, each with many objects.
  2. Run a pipeline that resolves s3://mybucket/sampleA/ on the head node (e.g. nf-schema input with format: path).
  3. Both prefixes are listed; the run hangs in submitted with an empty log.

Renaming to a lexicographically isolated prefix (e.g. sampleA-bc/) avoids the over-listing and the run starts normally.

Program output
Exception in thread "AwsEventLoop 7" java.lang.IllegalArgumentException: transferredBytes (481256976) must not be greater than totalBytes (481256866)
  at software.amazon.awssdk.transfer.s3.internal.progress.DefaultTransferProgressSnapshot.<init>(DefaultTransferProgressSnapshot.java:45)
  at software.amazon.awssdk.transfer.s3.internal.progress.TransferProgressUpdater.incrementBytesTransferred(TransferProgressUpdater.java:213)
  at software.amazon.awssdk.services.s3.internal.crt.S3CrtResponseHandlerAdapter.onProgress(S3CrtResponseHandlerAdapter.java:290)
Environment
  • Nextflow version: 25.10.5
  • Operating system: Linux (head job, Fusion-enabled CE)
Additional context

nf-amazon pins awssdk:s3 / s3-transfer-manager / aws-crt-client = 2.33.2. Likely fix: append / to the prefix and/or set .delimiter("/") as S3Iterator does; matchName() semantics are preserved. The CRT crash is upstream (aws/aws-sdk-java-v2#4790).

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start in S3ObjectSummaryLookup.lookup() and compare its request construction with S3Iterator.buildRequest(), which already applies a delimiter. Verify the behavior with sibling prefixes such as sampleA/ and sampleA-extra/. Done means lookup limits results to the requested directory while preserving matchName() semantics and avoiding the excess pagination described here.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws, groovy
Domain
cloud
Issue type
Bug
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Quiet
Clarity
Clearly specified
Newbie friendliness
70/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.