apache / apache/polaris

dropTable() API with Purge option as true is dropping the table in Polaris Rest catalog and deleted only the data files and not the metadata files in the storage(AWS S3)

Open
#1,448 2 comments 0 reactions 0 assignees View on GitHub
bug
Dominant language
Java
Stars
2.1k
Forks
522
Avg merge
2d 1h
Merged PRs (30d)
140

Description

### Describe the bug

We followed https://github.com/AlexMercedCoder/apache-polaris-learing-environment to bring up the Polaris catalog in one of our VM.

Created the catalogs, schemas and iceberg tables . As part of one of the operation we need to drop the table created in the Polaris catalog.

We are using a standalone SPARK application to try the options available where I have come across an issue with SPARK with Rest catalog.

Using the following code tried to drop the table

**Expectations :**
1. Drop the table from the Polaris catalog.
2. delete the metadata file from the storage (S3)
3. delete the data files from the storage (S3).

**Observations**
1.Dropped the table from the Polaris catalog.
2.deleted the data files from the storage (S3).

Metadata file from the storage is not deleted.

val catalog = spark.sessionState.catalogManager.catalog("dev_catalog").asInstanceOf[SparkCatalog]
val idnt = TableIdentifier.of("organization","finance")
catalog.icebergCatalog().dropTable(idnt,true)

### To Reproduce

Spark Application :

import org.apache.spark.sql.{ Row, Column, DataFrame, SaveMode, SparkSession ,Dataset}
//import software.amazon.awssdk.regions.Region
import scala.collection.mutable.HashMap
import org.apache.spark.sql.SparkSession
import org.apache.spark.sql.SparkSession
import org.apache.spark.SparkContext
import org.apache.spark.sql.SparkSession
import org.apache.spark.SparkContext
import org.apache.hadoop.conf.Configuration
import org.apache.hadoop.fs.FileSystem
import org.apache.hadoop.fs.Path
import java.net.URI
import org.apache.hadoop.conf.Configuration
import software.amazon.awssdk.services.sts.StsClient
import software.amazon.awssdk.services.sts.model.AssumeRoleRequest
import software.amazon.awssdk.services.sts.StsClient
import org.apache.iceberg.spark.SparkCatalog
import org.apache.iceberg.catalog.TableIdentifier;

import scala.Array
import org.apache.iceberg.rest.RESTCatalog
import org.apache.iceberg.spark.SparkCatalog

object ec2_check2_delete_stagingfile2{

def main(args: Array[String]): Unit = {
val spark = SparkSession
.builder()
.master("local[*]")
.config("spark.sql.extensions", "org.apache.iceberg.spark.extensions.IcebergSparkSessionExtensions")
.config("spark.sql.catalog.dev_catalog", "org.apache.iceberg.spark.SparkCatalog")
.config("spark.sql.catalog.dev_catalog.catalog-impl","org.apache.iceberg.rest.RESTCatalog")
.config("spark.sql.catalog.dev_catalog.uri","http://*************:8181/api/catalog")
.config("spark.sql.catalog.dev_catalog.header.X-Iceberg-Access-Delegation", true)
.config("spark.sql.catalog.dev_catalog.header.X-Iceberg-Access-Delegation","vended-credentials")
.config("spark.sql.catalog.dev_catalog.credential","*****:******")
.config("spark.sql.catalog.dev_catalog.client.region","******")
.config("spark.sql.catalog.dev_catalog.warehouse","dev_catalog")
.config("spark.sql.catalog.dev_catalog.scope","*****")
.config("spark.sql.catalog.dev_catalog.token-refresh-enabled", true)
.config("spark.sql.debug.codegen", true)
.getOrCreate();

print("Spark Running")

val catalog = spark.sessionState.catalogManager.catalog("dev_catalog").asInstanceOf[SparkCatalog]
val idnt = TableIdentifier.of("organization","finance")
catalog.icebergCatalog().dropTable(idnt,true)

print("done")

spark.stop();


}
}

### Actual Behavior

1.Dropped the table from the Polaris catalog.
2.deleted the data files from the storage (S3).

### Expected Behavior

1. Drop the table from the Polaris catalog.
2. delete the metadata file from the storage (S3)
3. delete the data files from the storage (S3).

### Additional context

_No response_

### System information

Dependencies :

1. iceberg-aws-bundle-1.4.3
2. iceberg-spark-runtime-3.3_2.12-1.4.3
3. log4j-slf4j-impl-2.17.2
4. iceberg-hive-runtime-1.6.1

Spark Version : 3.3.1
Scala Version : 2.12

Contributor guide

Open the contributing guide

Research direction

Start by reproducing the issue with the Spark application and the RESTCatalog/SparkCatalog dropTable(idnt, true) entry point described in the report. Compare the Polaris catalog state with the S3 contents after the operation, focusing on why data files are removed while metadata files remain. Done means the table, data files, and metadata files are all removed, with regression coverage added where the project’s tests support it.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws, java, scala
Domain
cloud, data-engineering, databases
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.