lpantano / lpantano/seqbuster

miraliger ignoring reads with length < 18 nt

Open
#24 5 comments 0 reactions 0 assignees View on GitHub
Dominant language
Java
Stars
14
Forks
3
PR merge metrics
No merged PRs in 30d

Description

Hi Lorena,

At first, thanks for a wonderful tool!

I have recently encountered an interesting 'feature' of the `miraligner` tool - it seems it's ignoring reads with less than 18 nucleotides. When I check both mapped (first column of .mirna) and unamapped reads (.mirna.nomap) there are not reads with less than 18 nucleotides. Is there some reason why miraligner ignores such reads? miRBase for human (for example) contains several mature miRNAs with less than 18 nucleotides and I would like to include them as well. I am aware that those miRNAs have a good potential to be included in the database just because some misannotation but still I would like to include such reads in my analysis.

Thank,
Jan

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by tracing how miraligner handles reads shorter than 18 nucleotides and how it writes the .mirna and .mirna.nomap outputs. Reproduce the behavior with short reads and verify that reads below 18 nt, including the relevant mature miRNAs, are represented in the appropriate output.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
bioinformatics
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.