aws / aws/sagemaker-python-sdk

Configurable (or just much longer?) health-check timeout in local mode

Offen
#3,362 2 Kommentare 2 Reaktionen 1 zugewiesene Person Beansprucht von @nargokul Auf GitHub ansehen
component: local mode type: feature request
Vorherrschende Sprache
Python
Sterne
2.3k
Forks
1.3k
Ø Merge
1 T. 22 Std.
Gemergte PRs (30 T.)
35

Beschreibung

**Describe the feature you'd like**

Today, local-mode endpoint deployment uses a [hard-coded health check time-out](https://github.com/aws/sagemaker-python-sdk/blob/ace07d72f4f44c43fe95b05574968decc7e806ac/src/sagemaker/local/entities.py#L44) of 120s for the container to become healthy.

This does not appear to be consistent with the start-up requirements for actual SageMaker endpoints, and even if it was, it may not be appropriate to assume local environments have similar network bandwidth or compute capabilities to target instance types.

**How would this feature be used? Please describe.**

I'm currently testing a use case with large (e.g. ~5GB+) model archives, and finding local mode deployment fails due to this healthcheck time-out, even though actual SageMaker endpoint deployments succeed without any issue.

If the default timeout was significantly longer, I think it should work okay. If the default timeout was configurable somehow, I could force it to wait longer for my use case.

**Describe alternatives you've considered**

Possible options could include:
- Extending the timeout
- Making the timeout configurable
- Somehow excluding tarball download/extract time from the coverage of the timeout check
- Supporting decompressed local folders as `model_data` targets for local models/endpoints - instead of requiring S3/tarball.

**Additional context**

N/A

Beitragsleitfaden

Beitragsleitfaden öffnen

Bewertung

Dieses Issue wurde noch nicht bewertet.

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.