aws / aws/sagemaker-huggingface-inference-toolkit
Using custom inference script and models from Hub
- Dominant language
- Python
- Stars
- 270
- Forks
- 60
- PR merge metrics
- No merged PRs in 30d
Description
Hello,
Can we use custom inference script without downloading the model to the S3 ? From the [link](https://huggingface.co/docs/sagemaker/inference#create-a-model-artifact-for-deployment), it says that we need to download the model artifacts and push them to S3 before using the custom inference script.
This will add considerable overhead depending on the model size.
Contributor guide
Research direction
The issue names no repository files, tests, or entry points. Start by reviewing the linked SageMaker custom-inference model-artifact requirements and tracing how this toolkit loads Hub models. Done means defining and validating a supported deployment path that avoids uploading model artifacts to S3.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- huggingface, python
- Domain
- cloud, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100