aws / aws/amazon-sagemaker-examples
[Example Request]Hosting multiple booster models using SageMaker Multi-model endpoints on NVIDIA Triton
- Dominant language
- Jupyter Notebook
- Stars
- 11k
- Forks
- 7k
- Avg merge
- 8h 29m
- Merged PRs (30d)
- 8
Description
**Describe the use case example you want to see**
Customers are looking to hosting multiple booster models using NVIDIA Triton and leverage SageMaker's multi-model endpoints to dynamically load and unload the models to save costs.
**How would this example be used? Please describe.**
This example demonstrates how customers can bring their pre-trained XGBoost models and host them on CPU and GPU on Triton using SageMaker multi-model endpoints. Also, this examples demos how MME dynamically loads/unloads the models.
**Describe which SageMaker services are involved**
SageMaker Hosting
**Describe what other services (other than SageMaker) are involved***
SageMaker, S3, ECR, IAM
**Describe which dataset could be used. Provide its location in s3://sagemaker-sample-files or another source.**
N/A
Contributor guide
Research direction
Review existing SageMaker Hosting, NVIDIA Triton, and multi-model endpoint examples before choosing the notebook entry point. Done means an example shows pre-trained XGBoost models hosted on CPU and GPU with SageMaker multi-model endpoints, including dynamic model loading and unloading and the stated S3, ECR, and IAM services.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- aws, jupyter-notebook
- Domain
- cloud, machine-learning
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100