Improve the split read tasks logic
Nobody has claimed this yet.
- Dominant language
- C++
- Stars
- 1k
- Forks
- 423
- Avg merge
- 1d 15h
- Merged PRs (30d)
- 24
Description
Enhancement
The current trySplitReadTasks cannot handle the case: if the expected stream num is less than segment num, the segment will become the basic unit for IO scheduling. And if there are some really large segments, the read data cannot be distributed evenly among the read threads which may cause the performance degrade.
What's more, we may also use region range as the basic IO scheduling unit in some other cases. And after dynamic region is supported, this may not continue to be a good choice. So we need to test it later after dynamic region is supported.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with trySplitReadTasks in dbms/src/Storages/DeltaMerge/SegmentReadTaskPool.cpp and examine how expected stream count and segment count determine IO scheduling. Define and validate a scheduling approach that distributes large-segment reads more evenly, while considering region ranges and the future dynamic-region behavior; the issue does not specify tests or a final design.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- cpp
- Domain
- databases, performance
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100