aws / aws/sagemaker-huggingface-inference-toolkit

Server reruns same task multiple times

Open
#133 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
270
Forks
60
PR merge metrics
No merged PRs in 30d

Description

I used
```
deploy = HuggingFaceModel(
name=model_name,
role=role,
code_location="abc",
model_data=path_to_s3,
transformers_version="4.37",
pytorch_version="2.1",
py_version='py310',
model_server_workers=1,
)
emb = deploy.deploy(
endpoint_name=model_name,
initial_instance_count=1,
instance_type="ml.c5.4xlarge",
container_startup_health_check_timeout=300,
)
```
The custom script was
```
def model_fn(model_dir):
processor = DataProcess() # A class that contains logic for processing each file.
return processor

def predict_fn(data, model):
text = model.process_file(data)
return {"output": text}
```
The input `data` is a base64 string of a file content.
It's strange that when the file is pretty small, under 1MB, the server runs `model_fn` and `predict_fn` once, and the process took around 30 seconds. But when I inputted large file of around 1.5MB, it runs `model_fn` and `predict_fn` multiple times, each time lasting around 2mins. I know this because the same request gives multiple contents of
```
[INFO ] W-model-1-stdout com.amazonaws.ml.mms.wlm.WorkerLifeCycle - Preprocess time - 5.128383636474609 ms
[INFO ] W-model-1-stdout com.amazonaws.ml.mms.wlm.WorkerLifeCycle - Predict time - 162199.17178153992 ms
[INFO ] W-model-1-stdout com.amazonaws.ml.mms.wlm.WorkerLifeCycle - Postprocess time - 0.00762939453125 ms
```

It's probably unorthodox to use the server for the data processing job. But what configs did I miss?

Related: https://github.com/aws/amazon-sagemaker-examples/issues/1073

Contributor guide

Open the contributing guide

Research direction

Start with the HuggingFaceModel deploy call and the model_fn and predict_fn entry points shown in the report. Reproduce the contrast between a sub-1MB and approximately 1.5MB base64 request while inspecting worker lifecycle logs, then determine whether request handling or a missing configuration causes repeated execution; done means the cause and applicable configuration are documented and the request runs once.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws, huggingface, python
Domain
backend, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.