[FEATURE] A more thorough data cleaning mechanism
- Dominant language
- Java
- Stars
- 454
- Forks
- 172
- Avg merge
- 5d 17h
- Merged PRs (30d)
- 5
Description
### Code of Conduct
- [X] I agree to follow this project's [Code of Conduct](https://www.apache.org/foundation/policies/conduct)
### Search before asking
- [X] I have searched in the [issues](https://github.com/apache/incubator-uniffle/issues?q=is%3Aissue) and found no similar issues.
### Describe the feature
We found that after the server cluster is restarted, the app directories on HDFS may not be cleaned up, which can lead to long-term occupation of disk space. We need a more thorough data cleaning mechanism to make sure data is always cleaned up, without leaks.
### Motivation
_No response_
### Describe the solution
_No response_
### Additional context
_No response_
### Are you willing to submit PR?
- [ ] Yes I am willing to submit a PR!
Contributor guide
Research direction
Start by tracing server-cluster restart handling and the lifecycle of app directories on HDFS. Define and verify behavior that removes leftover app directories after restarts without data leaks; the issue names no files or tests, so the relevant entry points must first be located.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- hadoop, java
- Domain
- data-engineering, distributed-systems
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100