aws / aws/sagemaker-python-sdk

Add output_data property to EstimatorBase class

Offen
#1,936 1 Kommentar 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
contributions welcome type: feature request
Vorherrschende Sprache
Python
Sterne
2.3k
Forks
1.3k
Ø Merge
1 T. 22 Std.
Gemergte PRs (30 T.)
35

Beschreibung

**Describe the feature you'd like**

As documented here:
https://sagemaker.readthedocs.io/en/stable/api/training/estimators.html#sagemaker.estimator.EstimatorBase.model_data, the `EstimatorBase` class provides a convenient method to pointing to the model .tar.gz archive location in S3 once `estimator.fit()` has been called.

Additionally to model data, SageMaker provides the ability to generate "output data" (different from "model data") when dumping files during training (e.g. experiment logs) to the directory defined by the environment variable `SM_OUTPUT_DATA_DIR`

Having a similar property in the `EstimatorBase` class, pointing to the output data .tar.gz archive location in S3 would be useful for developers wishing to manipulate that archive. It could be named, for example, `estimator.output_data`

**How would this feature be used? Please describe.**

Example use case:
1. Fit an estimator
2. During training, dump some files of interest to output data dir
3. Files are archived in output.tar.gz by SageMaker after .fit()
4. Access S3 location by using `estimator.output_data` and download output data archive locally.
5. Use output data archive locally, e.g. consult experiment logs

**Describe alternatives you've considered**

Compute the output_data location manually (potentially re-using the `.model_data` property)

Beitragsleitfaden

Beitragsleitfaden öffnen

Rechercherichtung

Beginne mit dem Lesen der model_data-Eigenschaft von EstimatorBase und des fit()-Ablaufs, der den S3-Speicherort von output.tar.gz bestimmt. Verfolge, wie SM_OUTPUT_DATA_DIR nach dem Training archiviert wird, und überprüfe anschließend, dass die neue output_data-Eigenschaft diesen Archivspeicherort nach Abschluss von fit() offenlegt.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
aws, python
Bereich
cloud, machine-learning
Issue-Typ
Feature
Schwierigkeit
3/5
Geschätzter Aufwand
1-2 Tage
Aktivitätsstatus
Veraltet
Klarheit
Größtenteils klar
Anfängerfreundlichkeit
35/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.