aws / aws/sagemaker-python-sdk

HyperparameterTuner does not preserve `disable_output_compression` setting of Estimator

未关闭
#5,206 2 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
component: training type: bug
主要语言
Python
星标
2.3k
派生
1.3k
平均合并
1 天 22 小时
30 天内合并 PR
35

描述

**Describe the bug**
When setting `disable_output_compression=True` in a PyTorch SageMaker Estimator that is used with HyperparameterTuner, the resulting HPO training jobs still compress the model output. This parameter is correctly set in the estimator but appears to be ignored when the estimator is used within a hyperparameter tuning job.

**To reproduce**
1. Create a PyTorch estimator with `disable_output_compression=True`
2. Use this estimator with a `HyperparameterTuner`
3. Launch the hyperparameter tuning job
4. Observe that the model outputs in the resulting training jobs are still compressed

Here's a minimal code example that reproduces the issue:

```python
import sagemaker
from sagemaker.pytorch.estimator import PyTorch
from sagemaker.tuner import HyperparameterTuner, ContinuousParameter

# Set up the PyTorch estimator with disable_output_compression=True
role = "arn:aws:iam::123456789012:role/SageMakerRole"
estimator = PyTorch(
disable_output_compression=True, # This should disable compression
entry_point="train.py",
source_dir="./code",
role=role,
instance_count=1,
instance_type="ml.m5.xlarge",
py_version="py3",
)

# Set up the hyperparameter tuner
hyperparameter_ranges = {
"learning-rate": ContinuousParameter(0.001, 0.1)
}

tuner = HyperparameterTuner(
estimator=estimator,
objective_metric_name="validation-accuracy",
hyperparameter_ranges=hyperparameter_ranges,
metric_definitions=[
{
"Name": "validation-accuracy",
"Regex": "validation accuracy: ([0-9\\.]+)"
}
],
max_jobs=2,
max_parallel_jobs=2
)

# Launch the tuning job
tuner.fit({"train": "s3://bucket/path/to/training/data"})
```

Expected behavior
When `disable_output_compression=True` is set in the PyTorch estimator, all training jobs created by the `HyperparameterTuner` should respect this setting and not compress the model output.

Screenshots or logs
When examining the training jobs created by the hyperparameter tuning job, the model artifacts are still compressed despite setting `disable_output_compression=True` in the estimator.

System information

* SageMaker Python SDK version: 2.246.0
* Framework name: PyTorch
* Framework version: 2.3.0
* Python version: 3.10
* CPU or GPU: CPU
* Custom Docker image: Yes, derived from `pytorch-training:2.3.0-cpu-py311-ubuntu20.04-sagemaker`

**Additional context**

This issue appears to be specific to hyperparameter tuning jobs. When using the same PyTorch estimator directly (not through a HyperparameterTuner), the disable_output_compression=True parameter works as expected and the model output is not compressed.

The issue might be that the disable_output_compression parameter is not being properly propagated from the estimator to the individual training jobs created by the HyperparameterTuner.

贡献指南

打开贡献指南

调研方向

首先跟踪 HyperparameterTuner 如何根据 PyTorch Estimator 构建单个训练作业的配置,然后将其与直接运行 estimator 时使用的配置进行比较。验证 disable_output_compression=True 在调优作业中得到保留,并通过在复现的场景中观察未压缩的模型构件来确认完成。

由索引模型根据 Issue 内容生成。

评估

技术栈
aws, python, pytorch
领域
cloud, machine-learning
Issue 类型
缺陷
难度
3/5
预计耗时
1-2 天
活跃度
停滞
描述清晰度
基本清楚
新手友好度
42/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。