Map common contaminants
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 21
- Forks
- 11
- Avg merge
- 12h 5m
- Merged PRs (30d)
- 24
Description
Add a few more references to the preliminary mapping step for things like phiX and E. coli.
Display a graph for each sample showing the relative amounts of target genomes and contaminants. Calculate the portions just based on read counts.
Make the remap step exclude any projects without coordinate references. That way we will calculate the portions of contaminants, but we won't carry them through the remapping and coverage steps.
Vera posted some example code on GitHub.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reviewing the preliminary mapping, remap, and coverage stages, then compare them with Vera's example code at github.com/veratai/check_miseq. Done means adding phiX and E. coli references, displaying per-sample read-count proportions, and excluding projects without coordinate references from remapping and coverage.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- bioinformatics, data
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 30/100