aws / aws/sagemaker-python-sdk

PySparkProcessor - Possibility to choose different instance types for the driver node and the worker nodes

Đang mở
#3,616 1 bình luận 0 reaction 1 người được giao Được @nargokul nhận Xem trên GitHub
component: processing PySpark
Ngôn ngữ chính
Python
Star
2.3k
Fork
1.3k
Merge trung bình
1 ngày 22 giờ
Pull request đã merge (30 ngày)
35

Mô tả

**Describe the feature you'd like**
It would be nice to have the possibility to choose different instance types for the driver node and the worker nodes when using the `PySparkProcessor`.

**How would this feature be used? Please describe.**
```
pyspark_processor = PySparkProcessor(
base_job_name=...,
framework_version=...,
role=...,
driver_instance_type="ml.m5.4xlarge",
worker_instance_type="ml.m5.large",
instance_count=...,
sagemaker_session=pipeline_session,
max_runtime_in_seconds=...,
)
```

**Describe alternatives you've considered**
It is possible to choose a high-memory instance for all instances but it could be unnecessarily costly for the user.

**Additional context**
Some pyspark operations (_e.g_. `.toPandas()`) are memory expensive for the driver node.

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.