Large spatial transcriptomics datasets best practices
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 394
- Forks
- 95
- Avg merge
- 4d 3h
- Merged PRs (30d)
- 7
Description
Hi spatialdata team,
Thank you for your work on this important project!
We are currently generating multiple terabytes of multi-modal digital spatial transcriptomics data, including H&E images, gene expression, spatial coordinates, and various annotations.
You can explore the initial 8 TB of data on HuggingFace, which contains over 56 million spots across 3,780 samples. Currently, each sample is stored as a separate compressed .h5ad.gz annata object.
We’d appreciate any guidance on best practices for storing and managing this scale of data using SpatialData. Have you benchmarked SpatialData for similarly large spatial transcriptomics datasets? Additionally, could you share what benefits we might expect from migrating our data to this format?
Thanks!
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the linked 8 TB HuggingFace dataset and its per-sample compressed .h5ad.gz layout. Investigate SpatialData's documented handling of multi-terabyte, multimodal spatial transcriptomics data and whether comparable benchmarks exist. Done means publishing concrete storage and management guidance, benchmark results, and clearly stated migration benefits.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- bioinformatics, data, documentation
- Issue type
- Documentation
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100