AbsaOSS / AbsaOSS/enceladus

Add ability to configure how Spark handles dates in parquet files.

Đang mở
#2,175 3 bình luận 0 reaction 1 người được giao Được @TebaleloS nhận Xem trên GitHub
Conformance feature priority: medium run scripts Standardization under discussion
Ngôn ngữ chính
Scala
Star
33
Fork
16
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Mô tả

## Background
With Spark 3 new option were added how to work with dates pre 1900 in parquet files
The settings are:
`spark.sql.parquet.datetimeRebaseModeInRead`
`spark.sql.parquet.datetimeRebaseModeInWrite`
`spark.sql.parquet.int96RebaseModeInRead`
`spark.sql.parquet.int96RebaseModeInWrite`

[Details here](https://spark.apache.org/docs/3.2.0/sql-data-sources-parquet.html).

## Feature
Allow setting of the options for Enceladus jobs
```[tasklist]
### Tasks
- [ ] ~Add command line options to be able to set the **read** options. Set a default behavior either to `EXCEPTION` or `LEGACY`.~
- [ ] ~Modify the helper scripts to recognize these settings~
- [ ] ~Add an `reference.conf`/`application.conf` setting to be applied to write options. The default should be `LEGACY`~
- [ ] Modify the helper scripts to be able to easily send the Spark settings into the `spark submit` - the defaults remain the same as described above
```

## To discuss
* The command line option names
* The command line defaults
* The write configuration names

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.