apache / apache/datafusion-comet

Enable spark.comet.exec.localTableScan.enabled when running Spark SQL tests

Open
#4,347 0 comments 0 reactions 0 assignees View on GitHub
priority:low spark sql tests
Dominant language
Scala
Stars
1.3k
Forks
373
Avg merge
2d 4h
Merged PRs (30d)
198

Description

Related to #2723.

A lot of Spark SQL tests don't put their source data in a format that Comet accelerates reading from. Comet generally wants data in Parquet or Iceberg. A number of Spark SQL suites (_e.g._, UDFSuite) don't write to Parquet so their scans are just `LocalTableScanExec`, and Comet is not exercising UDF compatibility because the scan underneath isn't native. We can either do what #2723 suggests and enable converting to columnar, or try enabling `LocalTableScanExec` now that we have #2735. @andygrove tried the former in #2714, but the logs are gone now so I'm not sure how ugly it was. Starting with `LocalTableScanExec` might be a smaller blast radius?

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.