Spark date handling settings can mess up dates in Standardization-Conformance combined job
- 主要言語
- Scala
- スター
- 33
- フォーク
- 16
- PR マージ指標
- 30日以内にマージされた PR はありません
説明
## Describe the bug
While #2175/#2184 solves for the ability to process dates prior to 1900 in Enceladus 3 (without code change), the way how it is done actually introduced a possible bug of messing up the dates.
In case a pair of different settings (LEGACY-CORRECTED) is used for read and write a combined job of Standardization&Conformance will mess up the dates.
## To Reproduce
Steps to reproduce the behavior OR commands run:
1. Have a dataset with timestamps pre 1900 from ambiguous interval
2. Run Standardization & Conformance job with settings read LEGACY, write CORRECTED
3. Dates will be messed up
## Expected behavior
Data should remain correct
## Additional context
Consider adding the information of used date reading standard into the _INFO file.
コントリビューションガイド
調査の方向性
Start by reproducing the Standardization and Conformance combined job with pre-1900 timestamps and LEGACY read versus CORRECTED write settings. Trace how the date-reading and date-writing standards are applied, then verify that dates remain correct and assess whether the used reading standard can be recorded in the _INFO file.
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- scala, spark
- 領域
- data-engineering
- issue の種類
- バグ
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 活発さ
- 停滞
- 明瞭さ
- おおむね明確
- 初心者へのやさしさ
- 35/100