aws / aws/sagemaker-huggingface-inference-toolkit

Endpoint creation completes before custom model_fn finishes loading resources

Open
#111 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
270
Forks
60
PR merge metrics
No merged PRs in 30d

Description

Image used: `763104351884.dkr.ecr.us-east-1.amazonaws.com/huggingface-pytorch-inference:2.1.0-transformers4.37.0-gpu-py310-cu118-ubuntu20.04` with custom `inference.py`.

I am loading some essential data from S3 for post processing on model's output and I included the data loading in the custom `model_fn` function. The data takes around 3~5 minutes to load. One thing I noticed is that the endpoint creation or update will complete before `model_fn` returns so the endpoint becomes available for incoming calls before all the data and model is loaded. This resulted in several minutes of additional latency around the period of time when endpoint is created or updated. How can I prevent this from happening?

Contributor guide

Open the contributing guide

Research direction

Start by tracing endpoint creation and update handling in the toolkit, then compare that lifecycle with the custom inference.py model_fn that loads resources from S3. Reproduce the delayed 3–5 minute load and observe when the endpoint begins accepting requests. Done should identify the relevant lifecycle behavior and document or implement a verified way to prevent requests before model_fn finishes.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws, python
Domain
backend, cloud
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.