AbsaOSS / AbsaOSS/enceladus

Catalyst optimizer throws an 'implicit cross join' exception when joining on a literal

Open
#947 0 comments 1 reaction 0 assignees View on GitHub
3rd party issue bug Conformance migration
Dominant language
Scala
Stars
33
Forks
16
PR merge metrics
No merged PRs in 30d

Description

## Describe the bug
The issue #895 is related to a bug in Catalyst optimizer resulting in the creation of an implicit cross join. This bug is known and there are 2 Spark jiras raised about this already.

## To Reproduce
The shortest Spark job to replicate the issue is described in [SPARK-29176](https://issues.apache.org/jira/browse/SPARK-29176):
```scala
case class Value(id: Int, lower: String, upper: String)

import spark.implicits._

val values = Seq(Value(1, "one", "ONE")).toDS
val join = values.join(values.withColumn("id", lit(1)), "id")
```

## Expected behaviour
If the original execution plan passes the implicit cross join test so should the optimized plan.

## Spark versions affected
- Spark 2.2.2 - no issue
- Spark 2.4.3, 2.4.4, master - the issue occurs

## Spark JIRAs to track
- [SPARK-29176](https://issues.apache.org/jira/browse/SPARK-29176)
- [SPARK-24839](https://issues.apache.org/jira/browse/SPARK-24839)

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.