AbsaOSS / AbsaOSS/enceladus

Optimise or allow for broadcast

Ouverte
#70 0 commentaires 0 réactions 1 personne assignée Réclamée par @yruslan Voir sur GitHub
Conformance feature priority: medium under discussion
Langage dominant
Scala
Étoiles
33
Forks
16
Métriques de merge des PR
Aucune PR mergée en 30 j

Description

"Automatic broadcast on join tables <10MB happens in spark 2.x + already, this covers a lot of the tables already"
References:
- https://stackoverflow.com/questions/43984068/does-spark-sql-autobroadcastjointhreshold-work-for-joins-using-datasets-join-op
- https://jaceklaskowski.gitbooks.io/mastering-spark-sql/spark-sql-joins-broadcast.html

In short:

> "Spark SQL uses broadcast join instead of hash join to optimize join queries when the size of one side data is below spark.sql.autoBroadcastJoinThreshold." "Spark will use autoBroadcastJoinThreshold and automatically broadcast data"

> "Broadcasting large objects is unlikely provide any performance boost, and in practice will often degrade performance and result in stability issue. Remember that broadcasted object has to be first fetch to driver, then send to each worker, and finally loaded into memory."

Guide de contribution

Ouvrir le guide de contribution

Évaluation

Cette issue n'a pas encore été évaluée.

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.