aws / aws/sagemaker-python-sdk
Enable processing and/or memory optimized instances when using sagemaker.remote_function's @remote decorator
- Langage dominant
- Python
- Étoiles
- 2.3k
- Forks
- 1.3k
- Merge moyen
- 1 j 22 h
- PR mergées (30 j)
- 35
Description
**Describe the feature you'd like**
Currently only training instances are allowed when using @remote (from the sagemaker.remote_function module). This module can be used for processing tasks as well, so it would be useful to have more instance types available.
**How would this feature be used? Please describe.**
Using instances with more than 256GB that don't need GPU acceleration for processing tasks. These are only available as Processing instances as far as I know (and are referred as Memory optimized instances there).
**Describe alternatives you've considered**
We can use Sagemaker Processing jobs, and we currently do that. The downside is that local mode is not enabled when using Sagemaker studio, so it can be a little clunky to develop scripts locally before submitting then to a processing task. This is much easier when using @remote, since we can execute code directly without the need of mapping inputs/outputs, etc. Code for local testing and remote execution could be very similar if not identical in this case.
**Additional context**
In case this is not clear, I'm referring to this functionality: https://docs.aws.amazon.com/sagemaker/latest/dg/train-remote-decorator.html
Link to the instance types available here: https://aws.amazon.com/sagemaker/pricing/
@jmahlik
Guide de contribution
Ouvrir le guide de contribution
Piste de recherche
Commencez par le module sagemaker.remote_function et son décorateur @remote, puis comparez les types d’instances acceptés avec la documentation SageMaker du décorateur remote indiquée dans l’issue. La tâche est terminée lorsque @remote prend en charge les instances de traitement et optimisées pour la mémoire pour les tâches de traitement, tout en préservant le comportement d’entraînement existant ; ajoutez une couverture ciblée si le dépôt contient des tests pertinents.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Évaluation
- Stack technique
- aws, machine-learning, python
- Domaine
- cloud, machine-learning
- Type d'issue
- Fonctionnalité
- Difficulté
- 4/5
- Temps estimé
- 3-5 jours
- Activité
- À l'abandon
- Clarté
- Plutôt claire
- Accessibilité débutants
- 38/100