galaxyproject / galaxyproject/idc

curation of the data

Open
#4 3 comments 0 reactions 0 assignees View on GitHub
Dominant language
Shell
Stars
10
Forks
8
PR merge metrics
No merged PRs in 30d

Description

Hi there. I'm currently discussing with our cluster admins if we can integrate the Galaxy data cache cvmfs.

They had a few questions:

1. is there a (defined) curation process of the data and maybe if there is some form of metadata for the data sets? I'm wondering in particular if its possible to determine the source of the data (e.g. if a genome was downloaded from NCBI/UCSC/..., if its with/without the mitogenome and other contigs, which data manager version was used for the creation, ... download date).

2. Are there methods implemented to ensure data integrity (eg checksums), in particular for data downloaded from public sources?

3. Is there a versioning system for the data? I have seen that there is eg for NCBI taxonomy, but genomes and indices seem not to be versioned.

And for my own curiosity: the files in this repo have surprisingly little content given the amount of data in http://datacache.galaxyproject.org/managed/.

Thanks...

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.