apache / apache/hudi

GCS incr source does not handle pubsub message properly

Open
#15,880 0 comments 0 reactions 0 assignees View on GitHub
area:ingest from-jira priority:high type:bug
Dominant language
Java
Stars
6.2k
Forks
2.5k
Avg merge
2d 8h
Merged PRs (30d)
111

Description

Gcs event source uses schema converter from spark and won't handle field name with hyphen in nested column. a sample message

{code:java}
23/04/03 19:23:45 DEBUG GcsEventsSource: msg: {
"kind": "storage#object",
"id": "",
"selfLink": "",
"name": "",
"bucket": "",
"generation": "1680505551370137",
"metageneration": "1",
"contentType": "application/octet-stream",
"timeCreated": "2023-04-03T07:05:51.373Z",
"updated": "2023-04-03T07:05:51.373Z",
"storageClass": "STANDARD",
"timeStorageClassUpdated": "2023-04-03T07:05:51.373Z",
"size": "6707",
"md5Hash": "",
"mediaLink": "",
"metadata": {
"goog-reserved-file-mtime": "1680503048"
},
"crc32c": "",
"etag": ""
}
{code}

and it throws

{code}
Exception in thread "main" org.apache.avro.SchemaParseException: Illegal character in: goog-reserved-file-mtime
at org.apache.avro.Schema.validateName(Schema.java:1571)
at org.apache.avro.Schema.access$400(Schema.java:92)
at org.apache.avro.Schema$Field.(Schema.java:549)
at org.apache.avro.SchemaBuilder$FieldBuilder.completeField(SchemaBuilder.java:2258)
at org.apache.avro.SchemaBuilder$FieldBuilder.completeField(SchemaBuilder.java:2254)
at org.apache.avro.SchemaBuilder$FieldBuilder.access$5100(SchemaBuilder.java:2150)
at org.apache.avro.SchemaBuilder$GenericDefault.noDefault(SchemaBuilder.java:2557)
at org.apache.hudi.org.apache.spark.sql.avro.SchemaConverters$.$anonfun$toAvroType$2(SchemaConverters.scala:205)
{code}

This is a problem with org.apache.spark.sql.avro.SchemaConverters#toAvroType

## JIRA info

- Link: https://issues.apache.org/jira/browse/HUDI-6028
- Type: Bug

Contributor guide

No contributing guide indexed for this repository

Research direction

Start at the GcsEventsSource entry point and trace how the Pub/Sub message reaches Spark's org.apache.spark.sql.avro.SchemaConverters#toAvroType. Reproduce the sample message with the hyphenated nested metadata field; done means the GCS incremental source processes it without the reported SchemaParseException.

Written by the indexing model from the issue text.

Assessment

Tech stack
google-cloud, java, spark
Domain
data-engineering, stream-processing
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.