aws / aws/sagemaker-huggingface-inference-toolkit
Num CPU incorrect
- Dominant language
- Python
- Stars
- 270
- Forks
- 60
- PR merge metrics
- No merged PRs in 30d
Description
Hi, deploying HF on TF image on ml.m5.4xlarge (16 vCPU)
I see in my launch logs `Number of CPUs: 1`
Is this still happening https://docs.aws.amazon.com/sagemaker/latest/dg/deploy-model-troubleshoot.html#deploy-model-troubleshoot-jvms ?
other related issue https://github.com/aws/sagemaker-inference-toolkit/issues/82
Contributor guide
Research direction
No source file or test is named. Start by reproducing the Hugging Face deployment on the stated SageMaker ml.m5.4xlarge TensorFlow image and inspect the launch logs, then compare the result with the linked SageMaker troubleshooting guidance and issue #82. Done means determining whether the reported CPU count is expected or identifying the inference-toolkit change needed to report the instance's 16 vCPUs.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- aws, python, tensorflow
- Domain
- backend, cloud, machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100