aws / aws/sagemaker-huggingface-inference-toolkit

Using custom inference script and models from Hub

Open
#102 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
270
Forks
60
PR merge metrics
No merged PRs in 30d

Description

Hello,

Can we use custom inference script without downloading the model to the S3 ? From the [link](https://huggingface.co/docs/sagemaker/inference#create-a-model-artifact-for-deployment), it says that we need to download the model artifacts and push them to S3 before using the custom inference script.

This will add considerable overhead depending on the model size.

Contributor guide

Open the contributing guide

Research direction

The issue names no repository files, tests, or entry points. Start by reviewing the linked SageMaker custom-inference model-artifact requirements and tracing how this toolkit loads Hub models. Done means defining and validating a supported deployment path that avoids uploading model artifacts to S3.

Written by the indexing model from the issue text.

Assessment

Tech stack
huggingface, python
Domain
cloud, machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.