aws / aws/sagemaker-python-sdk

Improved ways of storing local code in S3 for ProcessingSteps

未关闭
#4,879 0 条评论 0 个 reaction 已指派 1 人 已被 @mollyheamazon 认领 在 GitHub 查看
component: pipelines type: feature request
主要语言
Python
星标
2.3k
派生
1.3k
平均合并
1 天 22 小时
30 天内合并 PR
35

描述

**Describe the feature you'd like**
Currently, when using Processors such as `SKLearnProcessor` there is no way to specify where a local `code=` file should be stored in S3 when used in conjunction with a `ProcessingStep`. This can lead to clutter in S3 buckets, for example. The current behaviour places code in the `default_bucket` of a Sagemaker session like so:

`s3://{default_bucket}/auto_generated_hash/input/code/preprocess.py`

A better user experience would be to allow the user to define exactly where the code should be uploaded. This allows users to group files together for each run. For example:

`s3://{specified_bucket}/{project_name}/PIPELINE_EXECUTION_ID/code/preprocess.py`
`s3://{specified_bucket}/{project_name}/PIPELINE_EXECUTION_ID/data/train.csv`
`s3://{specified_bucket}/{project_name}/PIPELINE_EXECUTION_ID/model/model.pkl`

This should already be possible with the `FrameworkProcessor` and utilising the `code_location=` parameter but this seems to be ignored by the `ProcessingStep`.

贡献指南

打开贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。