apache / apache/gravitino

[EPIC] Make Gravitino run with cloud storage

Open
#4,396 0 comments 0 reactions 0 assignees View on GitHub
epic
Dominant language
Java
Stars
3.2k
Forks
935
Avg merge
1d 15h
Merged PRs (30d)
315

Description

### Describe the proposal

This EPIC issue aims to make Gravitino support running with cloud storage, like S3, ADLS, GCS, etc.

The background is that current Gravitino only tests against HDFS, but with the more demands of cloud object storage, we should make sure that Gravitino can work well with cloud storage.

The work here includes necessary code changes, configuration supports, validations and tests.

### Task list

- [ ] Hive catalog supports cloud storage.
- [ ] Iceberg catalog supports cloud storage
- [ ] Fileset supports cloud storage
- [ ] Iceberg rest catalog server supports cloud storage
- [ ] Paimon supports cloud storage
- [ ] Spark supports running on cloud storage
- [ ] Flink supports running on cloud storage
- [ ] Trino supports running on cloud storage

Contributor guide

Open the contributing guide

Research direction

The issue provides a broad task list covering Hive, Iceberg, Fileset, Paimon, Spark, Flink, and Trino, but names no files or tests. Start by selecting one unchecked integration and tracing its current HDFS configuration and validation tests. Done means that the selected component supports the target cloud storage, with configuration, validation, and tests updated.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws, azure, gcp, java
Domain
cloud, data-engineering
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.