Streaming Conformance too slow when just 20 rules are used
- Ngôn ngữ chính
- Scala
- Star
- 33
- Fork
- 16
- Chỉ số merge pull request
- Không có pull request nào được merge trong 30 ngày
Mô tả
## Describe the bug
Streaming Conformance takes way too much time to warm up (20 minutes) and to process. a single micro-batch (8 minutes).
This is way too slow.
This happens due to the catalyst issue reported to Spark:
https://issues.apache.org/jira/browse/SPARK-28090
And we have a workaround for batch: #190, #413
The job hangs completely if no workarounds are used: #1306
## To Reproduce
Steps to reproduce the behavior OR commands run:
1. Create a schema with nested arrays of structs.
2. Create a dataset with 20 conformance rules or so. Some of the rules should operate inside an array.
3. Run streaming conformance.
## Expected behaviour
Since all transformations are just projections the performance of a conformance job should be good.
## Additional context
Ideas of a solution:
- [ ] Tweak Catalyst workaround
- [ ] Rewrite conformance interpreter so that it uses only one single traversal for all conformance rules. We need to take in to account that rules are dependent. All dependencies need to be resolved in order to be applicable in a single traversal.
- [ ] Rewrite conformance interpreter using RDDs of `Row` + schema. Use a mapping lambda function to do conformance transformations in an imperative way. This way we use Spark only as a computation engine. Since no Catalyst is used, the Catalyst bug won't affect the computation.
Hướng dẫn đóng góp
Đánh giá
Issue này chưa được đánh giá.