apache / apache/hudi

df operation in Append mode for a MANAGED table results in "location already exists" error

Open
#13,931 2 comments 0 reactions 0 assignees View on GitHub
type:bug
Dominant language
Java
Stars
6.2k
Forks
2.5k
Avg merge
2d 8h
Merged PRs (30d)
111

Description

### Bug Description

**What happened:**
Running df.saveAsTable operations in Append mode for a MANAGED table results in `"Can not create the managed table('$catalogTableName') The associated location('$tableLocation') already exists".`
https://github.com/apache/hudi/blob/master/hudi-spark-datasource/hudi-spark-common/src/main/scala/org/apache/spark/sql/catalyst/catalog/HoodieCatalogTable.scala#L280-L281

**What you expected:**
Append mode should work for managed tables.

**Steps to reproduce:**
```
import org.apache.hudi.DataSourceWriteOptions
import org.apache.spark.sql.SaveMode

val df1 = Seq(
("100", "2015-01-01", "event_name_900", "2015-01-01T13:51:39.340396Z", "type1"),
("101", "2015-01-01", "event_name_546", "2015-01-01T12:14:58.597216Z", "type2")
).toDF("event_id", "event_date", "event_name", "event_ts", "event_type")

val tableName = ""
val databaseName = ""

df1.write.format("hudi")
.option("hoodie.metadata.enable", "true")
.option("hoodie.table.name", tableName)
.option("hoodie.database.name", databaseName)
.option("hoodie.datasource.write.operation", "upsert")
.option("hoodie.datasource.write.table.type", "COPY_ON_WRITE")
.option("hoodie.datasource.write.recordkey.field", "event_id")
.option("hoodie.datasource.write.precombine.field", "event_ts")
.option("hoodie.datasource.write.keygenerator.class", "org.apache.hudi.keygen.NonpartitionedKeyGenerator")
.option("hoodie.datasource.hive_sync.enable", "true")
.option("hoodie.datasource.meta.sync.enable", "true")
.option("hoodie.datasource.hive_sync.mode", "hms")
.option("hoodie.datasource.hive_sync.database", databaseName)
.option("hoodie.datasource.hive_sync.table", tableName)
.mode(SaveMode.Append)
.saveAsTable(s"$databaseName.$tableName")
```

### Environment

**Hudi version:** 1.0.2
**Query engine:** (Spark/Flink/Trino etc) Spark
**Relevant configs:**

### Logs and Stack Trace

`org.apache.spark.sql.AnalysisException: Can not create the managed table('spark_catalog..'). The associated location('/') already exists.`

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reproducing the Spark df.saveAsTable call in Append mode for a MANAGED table using the configuration and sample data in the issue. Then read hudi-spark-datasource/hudi-spark-common/src/main/scala/org/apache/spark/sql/catalyst/catalog/HoodieCatalogTable.scala around lines 280-281; done means the append completes without the managed-table location-already-exists error.

Written by the indexing model from the issue text.

Assessment

Tech stack
scala
Domain
data-engineering, databases
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.