apache / apache/sedona

Sedona 3.5 incompatible with cloudera jar versions?

Open
#3,097 4 comments 0 reactions 0 assignees View on GitHub
Dominant language
Java
Stars
2.4k
Forks
784
Avg merge
1d 12h
Merged PRs (30d)
58

Description

hi everyone!

I am getting an error when running a spark pipeline with sedona 1.7.2(also with 1.8.0 and 1.6.1) on Spark 3.5.4.3.5.7191000.100-2(cloudera environment):

```
java.lang.NoSuchMethodError: org.apache.spark.sql.execution.datasources.parquet.ParquetFooterReader.readFooter(Lorg/apache/hadoop/conf/Configuration;Lorg/apache/spark/sql/execution/datasources/PartitionedFile;Z)Lorg/apache/parquet/hadoop/metadata/ParquetMetadata;
at org.apache.spark.sql.execution.datasources.parquet.GeoParquetFileFormat.$anonfun$buildReaderWithPartitionValues$2(GeoParquetFileFormat.scala:240)
at org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.org$apache$spark$sql$execution$datasources$FileScanRDD$$anon$$readCurrentFile(FileScanRDD.scala:219)
at org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.nextIterator(FileScanRDD.scala:282)
at org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.hasNext(FileScanRDD.scala:131)
at scala.collection.Iterator$$anon$10.hasNext(Iterator.scala:460)
at org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIteratorForCodegenStage1.processNext(Unknown Source)
at org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)
at org.apache.spark.sql.execution.WholeStageCodegenEvaluatorFactory$WholeStageCodegenPartitionEvaluator$$anon$1.hasNext(WholeStageCodegenEvaluatorFactory.scala:43)
at org.apache.spark.sql.execution.UnsafeExternalRowSorter.sort(UnsafeExternalRowSorter.java:225)
at org.apache.spark.sql.execution.exchange.ShuffleExchangeExec$.$anonfun$prepareShuffleDependency$10(ShuffleExchangeExec.scala:377)
at org.apache.spark.rdd.RDD.$anonfun$mapPartitionsInternal$2(RDD.scala:922)
at org.apache.spark.rdd.RDD.$anonfun$mapPartitionsInternal$2$adapted(RDD.scala:922)
```

I have inspected the stack trace from:
```
val footerFileMetaData =
ParquetFooterReader
.readFooter(sharedConf, file, ParquetFooterReader.SKIP_ROW_GROUPS)
.getFileMetaData
```

the only difference that I see between 3.5.4 spark and cloudera flavor of the ParquetFooterReader is that this part is missing from the cloudera flavour, which matches the signature of the exception:
```
public class ParquetFooterReader {

public static final boolean SKIP_ROW_GROUPS = true;
public static final boolean WITH_ROW_GROUPS = false;

/**
* Reads footer for the input Parquet file 'split'. If 'skipRowGroup' is true,
* this will skip reading the Parquet row group metadata.
*
* @param file a part (i.e. "block") of a single file that should be read
* @param configuration hadoop configuration of file
* @param skipRowGroup If true, skip reading row groups;
* if false, read row groups according to the file split range
*/
public static ParquetMetadata readFooter(
Configuration configuration,
PartitionedFile file,
boolean skipRowGroup) throws IOException {
long fileStart = file.start();
ParquetMetadataConverter.MetadataFilter filter;
if (skipRowGroup) {
filter = ParquetMetadataConverter.SKIP_ROW_GROUPS;
} else {
filter = HadoopReadOptions.builder(configuration, file.toPath())
.withRange(fileStart, fileStart + file.length())
.build()
.getMetadataFilter();
}
return readFooter(configuration, file.toPath(), filter);
}
```

Has anyone encountered this issue on Cloudera environment? without shading this specific class, does any version work with CDS spark 3.5?
is there a way to use the unshaded sedona jars without conflict on CDP?

Contributor guide

Open the contributing guide

Research direction

Start at GeoParquetFileFormat.scala:240 and compare the ParquetFooterReader signature available in the stated Cloudera Spark 3.5 environment with the call shown in the issue. Reproduce the failure with Sedona 1.7.2 and the provided Spark version, then establish which supported jar version or packaging approach avoids the NoSuchMethodError.

Written by the indexing model from the issue text.

Assessment

Tech stack
java, scala
Domain
distributed-systems
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
38/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.