AlexsLemonade / AlexsLemonade/refinebio

Labeling single cell rna-seq datasets

Open
#972 0 comments 0 reactions 0 assignees View on GitHub
backlog single cell
Dominant language
Python
Stars
135
Forks
21
PR merge metrics
No merged PRs in 30d

Description

### Context

Through my looking for datasets for the tximport subset experiment, we've seen that some of the datasets in refine.bio are single cell RNA-seq and that most of the time (some of the time?)

### Problem or idea

We apparently haven't distinguished bulk rna-seq from single cell rna-seq? (Please correct me on what has been done on this front). It will be necessary to be able to separate single cell from bulk rna-seq in the future.

Although a lot of the initial processing methods will likely be the same or similar for single cell RNA-seq data (salmon and tximport), we will probably need to have an extra normalization step for single cell RNA-seq data. Also, in the future, we will probably not want single cell RNA-seq data and bulk RNA-seq data to be smashed together, so there may need to be some kind of block on that. (Can't think of a reason why users would want or need to compare single to bulk rna-seq. It isn't biologically sensical).

### Solution or next step

1) Single cell rna-seq datasets that are currently in refine.bio should be labeled in some way as to differentiate them from bulk rna-seq data.

2) It should be determined what (if any) single cell rna-seq datasets have been excluded from refine.bio pipeline and whether those datasets should be kept in mind or brought back for when a single cell RNA-seq processing pipeline is established in the future.

### New Issue Checklist

- [x] The title is short and descriptive.
- [x] You have explained the context that led you to write this issue.
- [x] You have reported a problem or idea.
- [x] You have proposed a solution or next step.

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reviewing how refine.bio currently represents datasets and how the pipeline handles Salmon and tximport; the issue names no files or tests. Identify current single-cell datasets and any that were excluded from the pipeline, then define the labeling and future handling needed before implementation.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
bioinformatics
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.