AbsaOSS / AbsaOSS/pramen

Allow specifying Spark resources per Python-Py Transformer

Open
#36 0 comments 1 reaction 0 assignees View on GitHub
enhancement help wanted Pramen-Py
Dominant language
Scala
Stars
31
Forks
4
Avg merge
1d 10m
Merged PRs (30d)
4

Description

## Background

Currently, Python transformations run with Spark resource settings configured per environment.
This is good if the cluster has dynamic resource allocation enabled. But for clusters that don't have dynamic resource allocation enabled it is not possible to fine-tune resources for each transformer.

## Solution
Pramen-Py already supports specifying resources per transformer #30. Need to add the capability to Pramen.

## Example
```conf
{
name = "Python transformer"
type = "python_transformation"
python.class = "SomeTransformer"
schedule.type = "daily"

info.date.expr = "@runDate"

output.table = "transformed_python"

dependencies = [
{
tables = [ table1 ]
date.from = "@infoDate"
}
]

spark.config {
spark.executor.instances = 4
spark.executor.cores = 1
spark.executor.memory = "4g"
}
}
```
Generated YAML:
```yaml
run_transformers:
- info_date: 2022-02-18
output_table: transformed_python
name: SomeTransformer
spark_config:
spark.executor.cores: 1
spark.executor.instances: 4
spark.executor.memory: 4g
```

Docs on Spark config: https://spark.apache.org/docs/latest/configuration.html

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.