aws / aws/sagemaker-python-sdk

MLFlow E2E Example Notebook

Ouverte
#5,513 0 commentaires 0 réactions 0 personnes assignées Voir sur GitHub
Langage dominant
Python
Étoiles
2.3k
Forks
1.3k
Merge moyen
1 j 22 h
PR mergées (30 j)
35

Description

**PR**: https://github.com/aws/sagemaker-python-sdk/pull/5514

**Describe the feature you'd like**

Add a new example notebook demonstrating the complete end-to-end workflow from training a PyTorch model to deploying it for inference on SageMaker, with MLflow 3.x tracking and model registry integration.

**How would this feature be used? Please describe.**

Users who want to leverage MLflow with SageMaker V3 SDK currently lack a comprehensive example showing the full workflow. This notebook would demonstrate:

1. Connecting to SageMaker MLflow tracking server
2. Training a PyTorch model with ModelTrainer while logging metrics/params to MLflow
3. Registering the trained model to MLflow Model Registry
4. Deploying directly from MLflow registry using ModelBuilder
5. Testing the deployed endpoint

Example workflow:

```
# Train with MLflow logging
model_trainer = ModelTrainer(training_image=..., source_code=...)
model_trainer.train()

# Deploy from MLflow registry
model_builder = ModelBuilder(
model_metadata={"MLFLOW_MODEL_PATH": "models:/my-model/1", ...}
)
model_builder.build()
model_builder.deploy()

```
**Describe alternatives you've considered**

Existing notebooks cover training or inference separately, but none show the integrated MLflow workflow end-to-end.

**Additional context**

Target location: v3-examples/ml-ops-examples/
Notebook name: v3-mlflow-train-inference-e2e-example.ipynb
Prerequisites: SageMaker MLflow App (tracking server ARN)

Guide de contribution

Ouvrir le guide de contribution

Piste de recherche

Commencez dans v3-examples/ml-ops-examples/ et utilisez le prérequis de l’ARN du serveur de suivi SageMaker MLflow App. Créez v3-mlflow-train-inference-e2e-example.ipynb couvrant l’entraînement PyTorch avec la journalisation MLflow, l’enregistrement du modèle, le déploiement via ModelBuilder et le test de l’endpoint ; le travail est considéré comme terminé lorsque le workflow complet est démontré.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
jupyter-notebook, python, pytorch
Domaine
cloud, documentation, machine-learning
Type d'issue
Fonctionnalité
Difficulté
3/5
Temps estimé
1-2 jours
Activité
À l'abandon
Clarté
Clairement spécifiée
Accessibilité débutants
25/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.