extract heterozygotes from the vcf before converting to h5 files
- Dominant language
- R
- Stars
- 10
- Forks
- 9
- PR merge metrics
- No merged PRs in 30d
Description
instead of downstream in the counts subworkflow
Homozygotes aren't useful in allele-specific analyses, so we discard them in the _counts_ subworkflow. But discarding them upstream, even before running WASP, might significantly speed up execution of the pipeline. So are there any downsides to this?
Contributor guide
No contributing guide indexed for this repository
Research direction
Begin with the counts subworkflow and the step that converts VCF data to H5 files; inspect where homozygotes are currently discarded and how WASP fits in. Determine whether moving filtering upstream preserves the pipeline’s allele-specific results and quantify any execution benefit before deciding what change is appropriate.
Written by the indexing model from the issue text.
Assessment
- Domain
- bioinformatics
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 30/100