apache / apache/gravitino

[Subtask] Fileset connector for PyTorch

Open
#2,549 1 comment 0 reactions 0 assignees View on GitHub
subtask
Dominant language
Java
Stars
3.2k
Forks
935
Avg merge
1d 15h
Merged PRs (30d)
315

Description

### Describe the subtask

Develop a PyTorch connector to load the data in a Gravitino Fileset. We can use this article as a reference:

https://docs.snowflake.com/en/developer-guide/snowpark-ml/snowpark-ml-framework-connectors

### Parent issue

#2113

Contributor guide

Open the contributing guide

Research direction

Start by reading parent issue #2113 and the linked Snowflake connector article to understand the expected connector behavior. Define the PyTorch connector's scope around loading data from a Gravitino Fileset; completion means that Fileset data can be loaded through PyTorch.

Written by the indexing model from the issue text.

Assessment

Tech stack
pytorch
Domain
data-engineering, machine-learning
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.