AbsaOSS / AbsaOSS/enceladus

Optimise or allow for broadcast

Đang mở
#70 0 bình luận 0 reaction 1 người được giao Được @yruslan nhận Xem trên GitHub
Conformance feature priority: medium under discussion
Ngôn ngữ chính
Scala
Star
33
Fork
16
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Mô tả

"Automatic broadcast on join tables <10MB happens in spark 2.x + already, this covers a lot of the tables already"
References:
- https://stackoverflow.com/questions/43984068/does-spark-sql-autobroadcastjointhreshold-work-for-joins-using-datasets-join-op
- https://jaceklaskowski.gitbooks.io/mastering-spark-sql/spark-sql-joins-broadcast.html

In short:

> "Spark SQL uses broadcast join instead of hash join to optimize join queries when the size of one side data is below spark.sql.autoBroadcastJoinThreshold." "Spark will use autoBroadcastJoinThreshold and automatically broadcast data"

> "Broadcasting large objects is unlikely provide any performance boost, and in practice will often degrade performance and result in stability issue. Remember that broadcasted object has to be first fetch to driver, then send to each worker, and finally loaded into memory."

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.