apache / apache/gravitino

[Subtask] Support read s3 fileset in Daft io

Open
#9,262 0 comments 0 reactions 0 assignees View on GitHub
subtask
Dominant language
Java
Stars
3.2k
Forks
935
Avg merge
1d 16h
Merged PRs (30d)
298

Description

### Describe the subtask

As titled, Gravitino fileset catalog supports multiple storages (s3, azure, gcs, hdfs). This issue is to support s3 fileset catalog reading in Daft io.

### Parent issue

https://github.com/apache/gravitino/issues/9259

Contributor guide

Open the contributing guide

Research direction

Start with the Daft io fileset-catalog integration and read the parent issue #9259 for the broader requirements. Trace how fileset catalogs for existing storage types are read, then verify that an S3-backed fileset can be read successfully through Daft io.

Written by the indexing model from the issue text.

Assessment

Domain
cloud, data
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.