AlexsLemonade / AlexsLemonade/scRNA-seq_sandbox

Filtering defaults for 10X

Open
#26 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
R
Stars
1
Forks
0
PR merge metrics
No merged PRs in 30d

Description

Through working with the Tabula Muris data I found:

1) No filtering of the data: does not work for scran: too many samples need to be dropped because of negative size factors.

2) The default filtering used for smart-seq2 is too stringent. For example: ~17000 samples were filtered down to ~8000, and ~27000 genes were filtered down to ~2000 genes.

### Question to answer:
What's a good starting point for 10X filtering?

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by examining the filtering behavior described for Tabula Muris 10X data and comparing it with the existing smart-seq2 default. Check how the no-filtering case interacts with scran, then document a justified starting point for 10X filtering and the expected effect on samples and genes.

Written by the indexing model from the issue text.

Assessment

Tech stack
r
Domain
data
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.