AlexsLemonade / AlexsLemonade/sc-data-integration
Future thought: Explore disagreements between references
- Dominant language
- HTML
- Stars
- 0
- Forks
- 0
- PR merge metrics
- No merged PRs in 30d
Description
This plot seems to me to be possibly the most useful one here. I am particularly interested by just comparing the blueprint to HPCA data, which seem to actually disagree on the labels "common myeloid progenitor" and "granulocyte monocyte progenitor" for a significant portion of cells. Blueprint "wins" here, but I honestly don't know what to make of that: is it a better reference, or just better matched to this dataset?
I think it might be interesting to plot a matrix of comparison between the the combined data and each reference, ideally some kind of annotation of the relationships among labels: when two references disagree, is there a parent/child relationship between the labels? We'd have to dive further into the ontology to create such a plot, so I think it is out of scope for this notebook. We would want to spend some time sketching out designs before trying to actually generate the plot, so this is a long-term thought.
_Originally posted by @jashapiro in https://github.com/AlexsLemonade/sc-data-integration/pull/219#discussion_r1207092879_
Contributor guide
No contributing guide indexed for this repository
Research direction
Begin by reviewing the existing notebook and the plot comparing Blueprint with HPCA data. The proposed work is to sketch and then generate a matrix comparing combined data with each reference, including ontology relationships such as parent and child labels; done means an agreed design and a useful comparison plot.
Written by the indexing model from the issue text.
Assessment
- Domain
- data, data-visualization
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100