rrwick / rrwick/Badread

Documentation on post processing of simulated sequences

Open
#22 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

enhancement
Dominant language
Python
Stars
302
Forks
25
PR merge metrics
No merged PRs in 30d

Description

Could you add additional documentation on the recommendations for post processing of the simulated sequencing data.

Following simulation of sequencing data, fastq or fastq.gz files were filtered and trimmed using X,Y and Z.

For example, typical ONT data is basecalled and demultiplexed using Guppy or Bonito. If the user includes adapters and barcodes in the sequences, should they demultiplex with guppy to permit filtering of sequencing WITHOUT a barcode on either end. Or should the user rely on more generic methods such as NanoFilt and cutadapt/trimmomatic?

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

No documentation file, test, or entry point is named. Review the existing documentation and the post-processing workflow for simulated FASTQ data, then resolve and document the recommendation for Guppy or Bonito versus generic filtering and trimming tools, including adapter and barcode handling.

Written by the indexing model from the issue text.

Assessment

Domain
bioinformatics, documentation
Issue type
Documentation
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.