Inferring pseudobulk batch effects
- Dominant language
- R
- Stars
- 102
- Forks
- 11
- PR merge metrics
- No merged PRs in 30d
Description
Hi!
I've been analyzing some scRNAseq data and noticed that randomly splitting and comparing replicates of the same condition after pseudobulking yields many differentially abundant genes with very small adjusted p-values. I'm guessing this is likely due to technical variation between replicates.
Given that LEMUR aligns the data at the single cell level, I was wondering if any of LEMUR's transformations used to align the data could also be used to infer sample-level batch effects for pseudobulk analysis with tools like DESeq or edgeR. Perhaps one or more covariates that can be included in the design matrix for these tools?
Thank you for your time and advice!
Contributor guide
No contributing guide indexed for this repository
Research direction
The issue does not name any files, tests, or entry points. Start by reviewing LEMUR's single-cell alignment transformations and how pseudobulk analyses with DESeq or edgeR use design-matrix covariates; done would be a documented, validated way to infer sample-level batch effects, if one exists.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- r
- Domain
- data
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100