aws / aws/sagemaker-huggingface-inference-toolkit

Num CPU incorrect

Open
#3 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
270
Forks
60
PR merge metrics
No merged PRs in 30d

Description

Hi, deploying HF on TF image on ml.m5.4xlarge (16 vCPU)
I see in my launch logs `Number of CPUs: 1`

Is this still happening https://docs.aws.amazon.com/sagemaker/latest/dg/deploy-model-troubleshoot.html#deploy-model-troubleshoot-jvms ?

other related issue https://github.com/aws/sagemaker-inference-toolkit/issues/82

Contributor guide

Open the contributing guide

Research direction

No source file or test is named. Start by reproducing the Hugging Face deployment on the stated SageMaker ml.m5.4xlarge TensorFlow image and inspect the launch logs, then compare the result with the linked SageMaker troubleshooting guidance and issue #82. Done means determining whether the reported CPU count is expected or identifying the inference-toolkit change needed to report the instance's 16 vCPUs.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws, python, tensorflow
Domain
backend, cloud, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.