[Improvement] Enhanced the table lineage for input tables
- Dominant language
- Scala
- Stars
- 2.4k
- Forks
- 1k
- PR merge metrics
- No merged PRs in 30d
Description
### Code of Conduct
- [X] I agree to follow this project's [Code of Conduct](https://www.apache.org/foundation/policies/conduct)
### Search before asking
- [X] I have searched in the [issues](https://github.com/apache/kyuubi/issues?q=is%3Aissue) and found no similar issues.
### What would you like to be improved?
```
spark.sql("CREATE TABLE t1 (a string, b string) USING hive")
spark.sql("CREATE TABLE t2 (a string, b string) USING hive")
val ret0 = exectractLineage("select t1.a from t1 where t1.b in (select b from t2)")
assert(ret0 == Lineage(
List("default.t1", "default.t2"),
List(),
List(("a", Set("default.t1.a")))))
val ret1 = exectractLineage("select t1.a from t1 join t2")
assert(ret1 == Lineage(
List("default.t1", "default.t2"),
List(),
List(("a", Set("default.t1.a")))))
```
In actual scenarios, it is necessary to display all input tables, even if a table may not contribute to the output columns.
### How should we improve?
_No response_
### Are you willing to submit PR?
- [X] Yes. I would be willing to submit a PR with guidance from the Kyuubi community to improve.
- [ ] No. I cannot submit a PR at this time.
Contributor guide
Research direction
Start from the lineage extraction path exercised by exectractLineage in the Spark SQL examples. Reproduce the nested-query and join cases, then trace how input tables and output-column lineage are collected. Done means both examples report all referenced input tables, including tables that do not contribute to output columns.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- scala, spark, sql
- Domain
- data-engineering, databases
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100