utilizing concurrency
@Sentdex is already working on this.
Since Jul 3, 2026.
Assessment
This issue has not been assessed yet.
Description
With pipeline parallelism and especially tensor parallelism, a lot of throughput performance is being left on the table by not solving any task that could be broken down into multiple pieces and solved in parallel.
Want to come up with a good way to utilize this extra perf, probably with some sort of toggle for max concurrency (default 1) and let the model div up tasks this way.
- Dominant language
- Python
- Stars
- 334
- Forks
- 41
- PR merge metrics
- No merged PRs in 30d
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Sentdex/minion
-
Difficulty 4/5 3-5 days Newbie friendliness 35/100
-
message box queue Open
Difficulty 5/5 Over a week Newbie friendliness 35/100
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
bancolombia/sentinel#23 ·
-
test md OpenCI
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
-
integration:quickjs org:external priority:backlog topic:code-interpreter topic:middleware type:feature
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
langchain-ai/deepagents#6450 ·
-
bug client
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100