aws / aws/sagemaker-python-sdk

HyperparameterTuner does not preserve `disable_output_compression` setting of Estimator

オープン
#5,206 コメント 2 件 リアクション 0 件 担当者 0 名 GitHub で見る
component: training type: bug
主要言語
Python
スター
2.3k
フォーク
1.3k
平均マージ
1日 22時間
マージ済み PR(30日)
35

説明

**Describe the bug**
When setting `disable_output_compression=True` in a PyTorch SageMaker Estimator that is used with HyperparameterTuner, the resulting HPO training jobs still compress the model output. This parameter is correctly set in the estimator but appears to be ignored when the estimator is used within a hyperparameter tuning job.

**To reproduce**
1. Create a PyTorch estimator with `disable_output_compression=True`
2. Use this estimator with a `HyperparameterTuner`
3. Launch the hyperparameter tuning job
4. Observe that the model outputs in the resulting training jobs are still compressed

Here's a minimal code example that reproduces the issue:

```python
import sagemaker
from sagemaker.pytorch.estimator import PyTorch
from sagemaker.tuner import HyperparameterTuner, ContinuousParameter

# Set up the PyTorch estimator with disable_output_compression=True
role = "arn:aws:iam::123456789012:role/SageMakerRole"
estimator = PyTorch(
disable_output_compression=True, # This should disable compression
entry_point="train.py",
source_dir="./code",
role=role,
instance_count=1,
instance_type="ml.m5.xlarge",
py_version="py3",
)

# Set up the hyperparameter tuner
hyperparameter_ranges = {
"learning-rate": ContinuousParameter(0.001, 0.1)
}

tuner = HyperparameterTuner(
estimator=estimator,
objective_metric_name="validation-accuracy",
hyperparameter_ranges=hyperparameter_ranges,
metric_definitions=[
{
"Name": "validation-accuracy",
"Regex": "validation accuracy: ([0-9\\.]+)"
}
],
max_jobs=2,
max_parallel_jobs=2
)

# Launch the tuning job
tuner.fit({"train": "s3://bucket/path/to/training/data"})
```

Expected behavior
When `disable_output_compression=True` is set in the PyTorch estimator, all training jobs created by the `HyperparameterTuner` should respect this setting and not compress the model output.

Screenshots or logs
When examining the training jobs created by the hyperparameter tuning job, the model artifacts are still compressed despite setting `disable_output_compression=True` in the estimator.

System information

* SageMaker Python SDK version: 2.246.0
* Framework name: PyTorch
* Framework version: 2.3.0
* Python version: 3.10
* CPU or GPU: CPU
* Custom Docker image: Yes, derived from `pytorch-training:2.3.0-cpu-py311-ubuntu20.04-sagemaker`

**Additional context**

This issue appears to be specific to hyperparameter tuning jobs. When using the same PyTorch estimator directly (not through a HyperparameterTuner), the disable_output_compression=True parameter works as expected and the model output is not compressed.

The issue might be that the disable_output_compression parameter is not being properly propagated from the estimator to the individual training jobs created by the HyperparameterTuner.

コントリビューションガイド

コントリビューションガイドを開く

調査の方向性

まず、HyperparameterTuner が PyTorch Estimator から個々のトレーニングジョブの設定をどのように構築するかを追跡し、次に、estimator を直接実行した場合に使用される設定と比較します。disable_output_compression=True がチューニングジョブで保持されていることを確認し、再現したシナリオで圧縮されていないモデルアーティファクトを観測して完了を確認します。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
aws, python, pytorch
領域
cloud, machine-learning
issue の種類
バグ
難易度
3/5
見積もり時間
1〜2日
活発さ
停滞
明瞭さ
おおむね明確
初心者へのやさしさ
42/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。