Split into separate pipelines?
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 21
- Forks
- 11
- Avg merge
- 12h 5m
- Merged PRs (30d)
- 24
Description
@jeff-k and I have been discussing how to include the T-cell receptor report into the MiCall pipeline for #461 . Currently, the pipeline tries to report on all the organisms it knows about, and generates reports for anything it finds. We then have a final validation step that checks whether the reports match the projects from the sample sheet. (Did we find what we expected to find?)
Instead of trying to report on all possible organisms, we could write completely separate pipelines for each one, and only try to report on the expected organisms for each sample. To detect contamination, we could use a general-purpose search tool like Kraken2 or Blast to identify what's in a sample, and check that against the sample sheet.
One challenge is how to handle mixed samples, such as V3LOOP and HLA sharing a pair of tags. We talked about three options:
- Run all the reads through both pipelines. It's possible that some reads could be counted as both V3LOOP and HLA.
- Run all the reads through the selected pipeline when a sample has only one organism expected, use Kraken2 to split the reads when a sample has more than one organism expected.
- Always use Kraken2 to identify the reads that get run through the selected pipelines, even when there's only one organism expected.
To see how common it is to mix samples, I scanned all of our old sample sheets. I found a total of 41469 samples, with the following project labels:
- HCV: 10901 up to 190401_M04401_0134_000000000-C4JMF
- V3LOOP: 8094 up to 190401_M04401_0134_000000000-C4JMF
- MidHCV: 5993 up to 190401_M04401_0134_000000000-C4JMF
- HLA-B: 3843 up to 180112_M04401_0099_000000000-BDR5L
- HIV: 2861 up to 190125_M01841_0375_000000000-C2W86
- MiniRT: 2057 up to 170627_M01841_0312_000000000-B8KHV
- V3LOOP-2: 1557 up to 131101_M01841_0037_000000000-A5F9E
- RT: 1256 up to 170502_M04401_0069_000000000-B2TD3
- Unknown: 1069 up to 190308_M04401_0132_000000000-C2VY6
- HLA-B, V3LOOP: 943 up to 180112_M04401_0099_000000000-BDR5L
- PR-RT: 651 up to 170606_M04401_0075_000000000-B2V2D
- HCV, HLA-B: 338 up to 160422_M01841_0227_000000000-AM8E8
- MiniHCV: 312 up to 150923_M01841_0172_000000000-AJHP5
- INT: 261 up to 170606_M04401_0075_000000000-B2V2D
- HCV-NS5a: 160 up to 170505_M01841_0301_000000000-B2T9R
- RT, V3LOOP: 99 up to 170502_M04401_0069_000000000-B2TD3
- RT-ONLY: 96 up to 130711_M01841_0010_000000000-A3TCY
- PR-RT, V3LOOP-2: 96 up to 130731_M01841_0014_000000000-A3TD3
- RT3-only, V3LOOP-2: 96 up to 130916_M01841_0022_000000000-A3RCA
- INT, PR-RT: 96 up to 151222_M01841_0201_000000000-AJGVR
- HLA-B, PR-RT: 90 up to 141120_M01841_0091_000000000-AA4HU
- HLA-B, MiniHCV, V3LOOP: 77 up to 150415_M01841_0125_000000000-ACF01
- HLA-B, Unknown: 71 up to 170714_M04401_0083_000000000-B9MJP
- HLA-B, MidHCV: 67 up to 160422_M01841_0227_000000000-AM8E8
- HXB2: 48 up to 130802_M01841_0015_000000000-A3RVT
- HCV, HLA-B, V3LOOP: 44 up to 160422_M01841_0227_000000000-AM8E8
- Unknown, V3LOOP: 39 up to 170627_M04401_0080_000000000-B2V5N
- HLA-B, MiniHCV: 38 up to 150713_M01841_0145_000000000-AE95N
- MiniHCV, V3LOOP: 35 up to 150415_M01841_0125_000000000-ACF01
- INT, PR-RT, V3LOOP: 32 up to 141104_M01841_0087_000000000-A8CGD
- HIV, HLA-B: 32 up to 170315_M01841_0291_000000000-AY9BB
- HCV, PR-RT: 31 up to 150921_M01841_0171_000000000-AJF7H
- HCV, V3LOOP: 23 up to 171027_M04401_0092_000000000-BBB77
- MidHCV, V3LOOP: 19 up to 171027_M01841_0331_000000000-BBB6R
- HLA-B, MidHCV, V3LOOP: 17 up to 160422_M01841_0227_000000000-AM8E8
- INT, V3LOOP: 11 up to 160415_M01841_0226_000000000-AM7BL
- HLA-B, INT: 10 up to 160415_M01841_0226_000000000-AM7BL
- PR-RT, V3LOOP: 5 up to 140925_M01841_0077_000000000-A8CDH
- HCV, MidHCV: 1 up to 150923_M01841_0172_000000000-AJHP5
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reviewing the current MiCall pipeline and sample-sheet handling, then examine issue #461 for the T-cell receptor context. Before implementation, resolve which pipeline arrangement and contamination strategy should support mixed samples; done means an agreed design with defined behavior for single- and multi-organism samples.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- bioinformatics
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100