aws / aws/sagemaker-python-sdk

Configurable (or just much longer?) health-check timeout in local mode

Aperta
#3,362 2 commenti 2 reazioni 1 assegnatario Rivendicata da @nargokul Vedi su GitHub
component: local mode type: feature request
Lingua principale
Python
Stelle
2.3k
Fork
1.3k
Merge medio
1g 22h
PR unite (30g)
35

Descrizione

**Describe the feature you'd like**

Today, local-mode endpoint deployment uses a [hard-coded health check time-out](https://github.com/aws/sagemaker-python-sdk/blob/ace07d72f4f44c43fe95b05574968decc7e806ac/src/sagemaker/local/entities.py#L44) of 120s for the container to become healthy.

This does not appear to be consistent with the start-up requirements for actual SageMaker endpoints, and even if it was, it may not be appropriate to assume local environments have similar network bandwidth or compute capabilities to target instance types.

**How would this feature be used? Please describe.**

I'm currently testing a use case with large (e.g. ~5GB+) model archives, and finding local mode deployment fails due to this healthcheck time-out, even though actual SageMaker endpoint deployments succeed without any issue.

If the default timeout was significantly longer, I think it should work okay. If the default timeout was configurable somehow, I could force it to wait longer for my use case.

**Describe alternatives you've considered**

Possible options could include:
- Extending the timeout
- Making the timeout configurable
- Somehow excluding tarball download/extract time from the coverage of the timeout check
- Supporting decompressed local folders as `model_data` targets for local models/endpoints - instead of requiring S3/tarball.

**Additional context**

N/A

Guida per i contributori

Apri la guida per i contributori

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.