[SUPPORT] Hudi could override users' configurations
- Dominant language
- Java
- Stars
- 6.2k
- Forks
- 2.5k
- Avg merge
- 2d 8h
- Merged PRs (30d)
- 111
Description
We recently also met the issue https://github.com/apache/hudi/issues/9305, but with the different cause(we still use hudi 0.12).
The user set the configure `spark.sql.parquet.enableVectorizedReader` to false manually, and read a hive table and cache it. Given spark will analyze the plan firstly if it needs to be cached, so currently spark won't add `C2R` to that cached plan since vectorized reader is false. At currently, spark won't execute that plan since there's no action operator.
Then user tries to read a MOR read_optimized table and join that cached plan and get the result, as mor table will automatically update the `enableVectorizedReader` to true, actually that hive table is read as column batch, but the plan doesn't contain `C2R` to convert the batch to row, whereas the error occurs:

```java
ava.lang.ClassCastException: org.apache.spark.sql.vectorized.ColumnarBatch cannot be cast to org.apache.spark.sql.catalyst.InternalRow
at scala.collection.Iterator$$anon$10.next(Iterator.scala:461)
at org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIteratorForCodegenStage1.processNext(Unknown Source)
at org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)
at org.apache.spark.sql.execution.WholeStageCodegenExec$$anon$1.hasNext(WholeStageCodegenExec.scala:759)
at org.apache.spark.sql.execution.columnar.DefaultCachedBatchSerializer$$anon$1.hasNext(InMemoryRelation.scala:118)
at scala.collection.Iterator$$anon$10.hasNext(Iterator.scala:460)
at org.apache.spark.storage.memory.MemoryStore.putIterator(MemoryStore.scala:223)
at org.apache.spark.storage.memory.MemoryStore.putIteratorAsValues(MemoryStore.scala:302)
at org.apache.spark.storage.BlockManager.$anonfun$doPutIterator$1(BlockManager.scala:1481)
at
```
```scala
override def imbueConfigs(sqlContext: SQLContext): Unit = {
super.imbueConfigs(sqlContext)
sqlContext.sparkSession.sessionState.conf.setConfString("spark.sql.parquet.enableVectorizedReader", "true")
}
```
I see there's some modification in the master code, but I suspect this issue could still happen since we'd also modify it in `HoodieFileGroupReaderBasedParquetFileFormat`:
```scala
spark.conf.set("spark.sql.parquet.enableVectorizedReader", supportBatchResult)
```
Besides this issue, Is it suitable to set spark configures globally? No matter users set it or not, I actually see hudi could set many spark relate configures in `SparkConf`, most of them are related to parquet reader/writer. This could confuse users and make it hard for devs to find the cause.
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reading imbueConfigs in the reported configuration path and the spark.conf.set assignment in HoodieFileGroupReaderBasedParquetFileFormat. Reproduce the cached Hive-table and MOR read-optimized join scenario with vectorized reading disabled, then verify that user settings are not unexpectedly overridden and the ColumnarBatch/InternalRow failure no longer occurs.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- java, scala
- Domain
- data-engineering, databases
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100