[Bug] case namenode rpc qps too high when all paritiion data written collectively
- Dominant language
- Java
- Stars
- 454
- Forks
- 172
- Avg merge
- 5d 17h
- Merged PRs (30d)
- 5
Description
### Code of Conduct
- [X] I agree to follow this project's [Code of Conduct](https://www.apache.org/foundation/policies/conduct)
### Search before asking
- [X] I have searched in the [issues](https://github.com/apache/incubator-uniffle/issues?q=is%3Aissue) and found no similar issues.
### Describe the bug
such as one shuffleId have too many partitions, if use MEMORY_HDFS OR MEMORY_LOCALFILE_HDFS storage type, when the stage startup, all parition data may collective flush to hdfs in a short time, so case namenode create rpc qps too high.
my solution:
I have change flush strategy, increase write priority for existing files,reduce write priority ofy for new files to avoid namenode create rpc too high.
do you have any better suggestions to solve this problem?
### Affects Version(s)
0.6.0
### Uniffle Server Log Output
_No response_
### Uniffle Engine Log Output
_No response_
### Uniffle Server Configurations
_No response_
### Uniffle Engine Configurations
_No response_
### Additional context
_No response_
### Are you willing to submit PR?
- [ ] Yes I am willing to submit a PR!
Contributor guide
Research direction
The issue concerns MEMORY_HDFS and MEMORY_LOCALFILE_HDFS storage types and excessive NameNode create RPC QPS during stage startup, but it names no source files or tests. Start by tracing the flush strategy for these storage types and reproduce the collective write pattern. Done means preventing the RPC spike without breaking partition-data flushing; the thread contains proposed prioritization but no decided approach.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- java
- Domain
- data-engineering, distributed-systems
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100