aws / aws/sagemaker-python-sdk

ModelTrainer and HyperparameterTuner missing environment variables

已关闭
#5,613 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
主要语言
Python
星标
2.3k
派生
1.3k
平均合并
1 天 22 小时
30 天内合并 PR
35

描述

**PySDK Version**
- [ ] PySDK V2 (2.x)
- [x] PySDK V3 (3.x)

**Describe the bug**
Tuning job does not add environment variables to the training jobs it creates. No environment variables are set, despite the environment variables being defined in the ModelTrainer correctly. The HyperparameterTuner does not seem to correctly propagate them.

**To reproduce**
Inside of a `TuningStep` pipeline step, use a `ModelTrainer` and add environment variables via the `environment` argument. Then define a `HyperparameterTuner` and pass the `model_trainer` object to it. Do all this inside of a pipeline session, such that `.tune()` will return the Arguments for the `TuningStep`. The resulting JSON of the pipeline definition will not contain an `"Environment"` key with environment variables inside the `"TrainingJobDefinition"` key within the `"Arguments"` for the tuning step.
```
model = ModelTrainer(
source_code=source_code,
compute=compute,
networking=self.networking,
base_job_name=base_job_name,
training_image=self.image_uris["train"],
output_data_config=OutputDataConfig(
s3_output_path=Join(
on="/",
values=[
self.s3_uri_runtime,
ExecutionVariables.PIPELINE_EXECUTION_ID,
"02_mt_output",
],
),
kms_key_id=self.aws_params["kms_key_hub"],
),
stopping_condition=StoppingCondition(max_runtime_in_seconds=28800),
role=self.aws_params["exec_role"],
sagemaker_session=self.pipeline_session,
environment={
"RANDOM_STATE": "42",
**self.default_env_vars,
},
# tags=self.tags_special,
)

hyperparameter_tuner = HyperparameterTuner(
model_trainer=model,
base_tuning_job_name=base_job_name,
metric_definitions=metric_definitions,
objective_metric_name=self.hpt_params["objective_metric_name"],
objective_type=self.hpt_params["objective_type"],
hyperparameter_ranges=self.hpt_params["hyperparameter_ranges"],
max_jobs=self.hpt_params["max_jobs"],
strategy="Bayesian",
max_parallel_jobs=4,
# random_seed=self.pipeline_params["RandomState"], # this is another bug, for another issue.
random_seed=42,
tags=self.tags,
)
```

**Expected behavior**
The environment variables passed as an argument to the `ModelTrainer` should be set in the training jobs created by a tuning job that uses this model trainer.

**System information**
A description of your system. Please provide:
- **SageMaker Python SDK version**: 3.5.0

**Additional context**
This is a roadblock for us regarding a migration from v2 to v3.

贡献指南

打开贡献指南

调研方向

从复现中使用的 ModelTrainer、HyperparameterTuner 和 TuningStep 入口点开始,然后检查 model trainer 如何被序列化到 tuning step 的 TrainingJobDefinition 中。复现 pipeline JSON,并验证 Environment 键包含 ModelTrainer 变量,以使结果完整。

由索引模型根据 Issue 内容生成。

评估

技术栈
aws, python
领域
machine-learning
Issue 类型
缺陷
难度
3/5
预计耗时
1-2 天
活跃度
停滞
描述清晰度
基本清楚
新手友好度
52/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。