google / google/Xee

Long-running code results in `requests` `ChunkedEncodingError` exception (broken connection)

Open
#125 1 comment 0 reactions 1 assignee Claimed by @noahgolmant View on GitHub
question
Dominant language
Python
Stars
371
Forks
48
PR merge metrics
No merged PRs in 30d

Description

I have a script ingesting ~200 GB of landsat imagery with the current multi-threaded implementation (no Dataflow). Eventually, I always get an exception like:

```
requests.exceptions.ChunkedEncodingError: ('Connection broken: IncompleteRead(9186238 bytes read, 1299762 more expected)', IncompleteRead(9
186238 bytes read, 1299762 more expected))
```

This occurs in the [`common.robust_getitem`](https://github.com/google/Xee/blob/f05e82b751d54433099c89c9ce7d68c629c592dc/xee/ext.py#L463) call.

I've had some success in reducing the frequency of this exception by lowering the chunk size, so e.g. I can make it to ~150 GB instead of failing after 90, although hard to say if that improvement is reliable since it is non-deterministic.

I am not sure of the root cause of this-- it could be due to a multithreading/lock issue, or the server is prematurely closing the connection. Either way, the current code only applies the retry/backoff logic to `EEException`s. I've had success by retrying on any `Exception` rather than just `EEException` but that is not an ideal solution.

I'd imagine that we don't see this in Dataflow because it has its own worker retry logic?

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.