aws / aws/sagemaker-python-sdk

Add output_data property to EstimatorBase class

未关闭
#1,936 1 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
contributions welcome type: feature request
主要语言
Python
星标
2.3k
派生
1.3k
平均合并
1 天 22 小时
30 天内合并 PR
35

描述

**Describe the feature you'd like**

As documented here:
https://sagemaker.readthedocs.io/en/stable/api/training/estimators.html#sagemaker.estimator.EstimatorBase.model_data, the `EstimatorBase` class provides a convenient method to pointing to the model .tar.gz archive location in S3 once `estimator.fit()` has been called.

Additionally to model data, SageMaker provides the ability to generate "output data" (different from "model data") when dumping files during training (e.g. experiment logs) to the directory defined by the environment variable `SM_OUTPUT_DATA_DIR`

Having a similar property in the `EstimatorBase` class, pointing to the output data .tar.gz archive location in S3 would be useful for developers wishing to manipulate that archive. It could be named, for example, `estimator.output_data`

**How would this feature be used? Please describe.**

Example use case:
1. Fit an estimator
2. During training, dump some files of interest to output data dir
3. Files are archived in output.tar.gz by SageMaker after .fit()
4. Access S3 location by using `estimator.output_data` and download output data archive locally.
5. Use output data archive locally, e.g. consult experiment logs

**Describe alternatives you've considered**

Compute the output_data location manually (potentially re-using the `.model_data` property)

贡献指南

打开贡献指南

调研方向

首先阅读 EstimatorBase 的 model_data 属性以及确定 output.tar.gz S3 位置的 fit() 流程。跟踪训练后如何归档 SM_OUTPUT_DATA_DIR,然后验证新的 output_data 属性是否会在 fit() 完成后公开该归档位置。

由索引模型根据 Issue 内容生成。

评估

技术栈
aws, python
领域
cloud, machine-learning
Issue 类型
功能
难度
3/5
预计耗时
1-2 天
活跃度
停滞
描述清晰度
基本清楚
新手友好度
35/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。