Allow specifying Spark resources per Python-Py Transformer
- 主要語言
- Scala
- 星號
- 31
- 分支
- 4
- 平均合併
- 1 天 10 分鐘
- 30 天內合併 PR
- 4
描述
## Background
Currently, Python transformations run with Spark resource settings configured per environment.
This is good if the cluster has dynamic resource allocation enabled. But for clusters that don't have dynamic resource allocation enabled it is not possible to fine-tune resources for each transformer.
## Solution
Pramen-Py already supports specifying resources per transformer #30. Need to add the capability to Pramen.
## Example
```conf
{
name = "Python transformer"
type = "python_transformation"
python.class = "SomeTransformer"
schedule.type = "daily"
info.date.expr = "@runDate"
output.table = "transformed_python"
dependencies = [
{
tables = [ table1 ]
date.from = "@infoDate"
}
]
spark.config {
spark.executor.instances = 4
spark.executor.cores = 1
spark.executor.memory = "4g"
}
}
```
Generated YAML:
```yaml
run_transformers:
- info_date: 2022-02-18
output_table: transformed_python
name: SomeTransformer
spark_config:
spark.executor.cores: 1
spark.executor.instances: 4
spark.executor.memory: 4g
```
Docs on Spark config: https://spark.apache.org/docs/latest/configuration.html
貢獻指南
這個儲存庫沒有索引到貢獻指南
評估
這個 Issue 還沒有評估資料。