aws / aws/amazon-sagemaker-examples

[Example Request]Hosting multiple booster models using SageMaker Multi-model endpoints on NVIDIA Triton

Open
#3,771 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
11k
Forks
7k
Avg merge
8h 29m
Merged PRs (30d)
8

Description

**Describe the use case example you want to see**
Customers are looking to hosting multiple booster models using NVIDIA Triton and leverage SageMaker's multi-model endpoints to dynamically load and unload the models to save costs.

**How would this example be used? Please describe.**
This example demonstrates how customers can bring their pre-trained XGBoost models and host them on CPU and GPU on Triton using SageMaker multi-model endpoints. Also, this examples demos how MME dynamically loads/unloads the models.

**Describe which SageMaker services are involved**
SageMaker Hosting

**Describe what other services (other than SageMaker) are involved***
SageMaker, S3, ECR, IAM

**Describe which dataset could be used. Provide its location in s3://sagemaker-sample-files or another source.**
N/A

Contributor guide

Open the contributing guide

Research direction

Review existing SageMaker Hosting, NVIDIA Triton, and multi-model endpoint examples before choosing the notebook entry point. Done means an example shows pre-trained XGBoost models hosted on CPU and GPU with SageMaker multi-model endpoints, including dynamic model loading and unloading and the stated S3, ECR, and IAM services.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws, jupyter-notebook
Domain
cloud, machine-learning
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.