Multiprocess for sc.tl.pca, sc.pp.neighbors, sc.tl.leiden, and sc.tl.umap
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 2.6k
- Forks
- 779
- Avg merge
- 1d 4h
- Merged PRs (30d)
- 27
Description
- Additional function parameters / changed functionality / changed defaults?
- New analysis tool: A simple analysis tool you have been using and are missing in
sc.tools? - New plotting function: A kind of plot you would like to seein
sc.pl? - External tools: Do you know an existing package that should go into
sc.external.*? - Other?
Hello Scanpy,
I'm merging millions of cells to run the Scanpy, and the steps for constructing UMAP kill me. Is it possible or is there any plan to make these functions sc.tl.pca, sc.pp.neighbors, sc.tl.leiden, and sc.tl.umap support multiprocessing?
Thanks!
Best,
Yuanjian
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by locating the implementations and existing tests for sc.tl.pca, sc.pp.neighbors, sc.tl.leiden, and sc.tl.umap. Determine which operations dominate runtime on millions of cells and how multiprocessing should apply across these entry points; done should mean documented, tested multiprocessing support for the requested functions.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- bioinformatics, machine-learning, performance
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 30/100