aws / aws/sagemaker-python-sdk

Add output_data property to EstimatorBase class

Open
#1,936 1 comment 0 reactions 0 assignees View on GitHub
contributions welcome type: feature request
Dominant language
Python
Stars
2.3k
Forks
1.3k
Avg merge
1d 22h
Merged PRs (30d)
35

Description

**Describe the feature you'd like**

As documented here:
https://sagemaker.readthedocs.io/en/stable/api/training/estimators.html#sagemaker.estimator.EstimatorBase.model_data, the `EstimatorBase` class provides a convenient method to pointing to the model .tar.gz archive location in S3 once `estimator.fit()` has been called.

Additionally to model data, SageMaker provides the ability to generate "output data" (different from "model data") when dumping files during training (e.g. experiment logs) to the directory defined by the environment variable `SM_OUTPUT_DATA_DIR`

Having a similar property in the `EstimatorBase` class, pointing to the output data .tar.gz archive location in S3 would be useful for developers wishing to manipulate that archive. It could be named, for example, `estimator.output_data`

**How would this feature be used? Please describe.**

Example use case:
1. Fit an estimator
2. During training, dump some files of interest to output data dir
3. Files are archived in output.tar.gz by SageMaker after .fit()
4. Access S3 location by using `estimator.output_data` and download output data archive locally.
5. Use output data archive locally, e.g. consult experiment logs

**Describe alternatives you've considered**

Compute the output_data location manually (potentially re-using the `.model_data` property)

Contributor guide

Open the contributing guide

Research direction

Start by reading the EstimatorBase model_data property and the fit() flow that determines the output.tar.gz S3 location. Trace how SM_OUTPUT_DATA_DIR is archived after training, then verify that the new output_data property exposes that archive location after fit() completes.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws, python
Domain
cloud, machine-learning
Issue type
Feature
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.