apache / apache/gravitino

[Improvement] Management of Small Files and Data Acceleration in AI Scenarios

Open
#3,995 0 comments 0 reactions 0 assignees View on GitHub
improvement
Dominant language
Java
Stars
3.2k
Forks
935
Avg merge
1d 15h
Merged PRs (30d)
315

Description

### What would you like to be improved?

In the AI scenario, managing related files with FileSet may have a scenario with too many small files. How does Gravitino consider the need to merge small files? In addition, how to accelerate data? And is there a consideration for implementing cold storage for related data?

### How should we improve?

_No response_

Contributor guide

Open the contributing guide

Research direction

Start by reviewing how FileSet manages related files and how the project handles data acceleration or cold storage; the issue names no files, tests, or entry points. Before implementation, clarify the desired small-file merge, acceleration, and cold-storage behavior, then define completion criteria for those requirements.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
ai, data, performance
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.