aws / aws/amazon-sagemaker-feedback

GPU/CUDA for Serverless Inference

Open
#233 0 comments 0 reactions 0 assignees View on GitHub
feature-request
Dominant language
No language data
Stars
10
Forks
3
PR merge metrics
No merged PRs in 30d

Description

### Product Version

- [ ] Amazon SageMaker Studio Classic
- [ ] Amazon SageMaker Studio
- [x] It is not related to SageMaker Studio

### Product Category

Inference and Endpoints

### Description

Serverless Inference is an excellent offering for workloads with intermittent or unpredictable traffic, thanks to its automatic scaling and pay-per-request pricing model. However, the lack of GPU variations currently limits its applicability for ML use cases that benefit from GPUs and CUDA acceleration.

### Other Details

_No response_

Contributor guide

Open the contributing guide

Research direction

The issue names Amazon SageMaker Serverless Inference but provides no repository files, tests, or implementation entry point. Start by reviewing the Serverless Inference product area and determine the requirements for GPU and CUDA variants; done would require a concrete product decision or implementation scope for that support.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws
Domain
cloud, machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.