[EPIC] Make Gravitino run with cloud storage
- Dominant language
- Java
- Stars
- 3.2k
- Forks
- 935
- Avg merge
- 1d 15h
- Merged PRs (30d)
- 315
Description
### Describe the proposal
This EPIC issue aims to make Gravitino support running with cloud storage, like S3, ADLS, GCS, etc.
The background is that current Gravitino only tests against HDFS, but with the more demands of cloud object storage, we should make sure that Gravitino can work well with cloud storage.
The work here includes necessary code changes, configuration supports, validations and tests.
### Task list
- [ ] Hive catalog supports cloud storage.
- [ ] Iceberg catalog supports cloud storage
- [ ] Fileset supports cloud storage
- [ ] Iceberg rest catalog server supports cloud storage
- [ ] Paimon supports cloud storage
- [ ] Spark supports running on cloud storage
- [ ] Flink supports running on cloud storage
- [ ] Trino supports running on cloud storage
Contributor guide
Research direction
The issue provides a broad task list covering Hive, Iceberg, Fileset, Paimon, Spark, Flink, and Trino, but names no files or tests. Start by selecting one unchecked integration and tracing its current HDFS configuration and validation tests. Done means that the selected component supports the target cloud storage, with configuration, validation, and tests updated.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- aws, azure, gcp, java
- Domain
- cloud, data-engineering
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100