AbsaOSS / AbsaOSS/enceladus

Add ability to configure how Spark handles dates in parquet files.

Aperta
#2,175 3 commenti 0 reazioni 1 assegnatario Rivendicata da @TebaleloS Vedi su GitHub
Conformance feature priority: medium run scripts Standardization under discussion
Lingua principale
Scala
Stelle
33
Fork
16
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Descrizione

## Background
With Spark 3 new option were added how to work with dates pre 1900 in parquet files
The settings are:
`spark.sql.parquet.datetimeRebaseModeInRead`
`spark.sql.parquet.datetimeRebaseModeInWrite`
`spark.sql.parquet.int96RebaseModeInRead`
`spark.sql.parquet.int96RebaseModeInWrite`

[Details here](https://spark.apache.org/docs/3.2.0/sql-data-sources-parquet.html).

## Feature
Allow setting of the options for Enceladus jobs
```[tasklist]
### Tasks
- [ ] ~Add command line options to be able to set the **read** options. Set a default behavior either to `EXCEPTION` or `LEGACY`.~
- [ ] ~Modify the helper scripts to recognize these settings~
- [ ] ~Add an `reference.conf`/`application.conf` setting to be applied to write options. The default should be `LEGACY`~
- [ ] Modify the helper scripts to be able to easily send the Spark settings into the `spark submit` - the defaults remain the same as described above
```

## To discuss
* The command line option names
* The command line defaults
* The write configuration names

Guida per i contributori

Apri la guida per i contributori

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.