aws / aws/sagemaker-python-sdk

PySparkProcessor - Possibility to choose different instance types for the driver node and the worker nodes

未关闭
#3,616 1 条评论 0 个 reaction 已指派 1 人 已被 @nargokul 认领 在 GitHub 查看
component: processing PySpark
主要语言
Python
星标
2.3k
派生
1.3k
平均合并
1 天 22 小时
30 天内合并 PR
35

描述

**Describe the feature you'd like**
It would be nice to have the possibility to choose different instance types for the driver node and the worker nodes when using the `PySparkProcessor`.

**How would this feature be used? Please describe.**
```
pyspark_processor = PySparkProcessor(
base_job_name=...,
framework_version=...,
role=...,
driver_instance_type="ml.m5.4xlarge",
worker_instance_type="ml.m5.large",
instance_count=...,
sagemaker_session=pipeline_session,
max_runtime_in_seconds=...,
)
```

**Describe alternatives you've considered**
It is possible to choose a high-memory instance for all instances but it could be unnecessarily costly for the user.

**Additional context**
Some pyspark operations (_e.g_. `.toPandas()`) are memory expensive for the driver node.

贡献指南

打开贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。