google-deepmind / google-deepmind/alphafold
comparing ready-made db vs. gsutil based download
- Dominant language
- Python
- Stars
- 14.9k
- Forks
- 2.9k
- PR merge metrics
- No merged PRs in 30d
Description
Thank you very much for the various pre-made and custom-downloads of AF2 databases that you have made available to the research community.
Before using the `gsutil` based download of shards for species specific file download to my local machine, I wanted to make sure these 2 types of downloads are effectively the same when I execute download steps locally. I chose _Arabidopsis thaliana_ as a test case.
In my test, for the species, Arabidopsis thaliana, the **pre-made AF2 Database vs. what I downloaded using `gsutil`** gave **different numbers of downloaded files**. Details below.
## Pre-made AF2 DB download for _Arabidopsis thaliana_
Using link at https://alphafold.ebi.ac.uk/download for pre-made AF2 database for this species:
`% wget https://ftp.ebi.ac.uk/pub/databases/alphafold/latest/UP000006548_3702_ARATH_v4.tar
`
Downloaded file details:
```
% ls -lht
total 7615488
-rwxrwxrwx 1 aksrao staff 3.6G Oct 20 04:55 UP000006548_3702_ARATH_v4.tar
```
## Using gsutil to download shards for _Arabidopsis thaliana_
Using instructions at https://github.com/deepmind/alphafold/blob/main/afdb/README.md, set up local machine for Google Cloud, and then:
```
% gsutil -m -o "GSUtil:parallel_process_count=1" cp "gs://public-datasets-deepmind-alphafold-v4/proteomes/proteome-tax_id-3702-*_v4.tar" .
Copying gs://public-datasets-deepmind-alphafold-v4/proteomes/proteome-tax_id-3702-0_v4.tar...
Copying gs://public-datasets-deepmind-alphafold-v4/proteomes/proteome-tax_id-3702-10_v4.tar...
Copying gs://public-datasets-deepmind-alphafold-v4/proteomes/proteome-tax_id-3702-11_v4.tar...
Copying gs://public-datasets-deepmind-alphafold-v4/proteomes/proteome-tax_id-3702-12_v4.tar...
Copying gs://public-datasets-deepmind-alphafold-v4/proteomes/proteome-tax_id-3702-13_v4.tar...
Copying gs://public-datasets-deepmind-alphafold-v4/proteomes/proteome-tax_id-3702-1_v4.tar...
Copying gs://public-datasets-deepmind-alphafold-v4/proteomes/proteome-tax_id-3702-2_v4.tar...
Copying gs://public-datasets-deepmind-alphafold-v4/proteomes/proteome-tax_id-3702-3_v4.tar...
Copying gs://public-datasets-deepmind-alphafold-v4/proteomes/proteome-tax_id-3702-4_v4.tar...
Copying gs://public-datasets-deepmind-alphafold-v4/proteomes/proteome-tax_id-3702-5_v4.tar...
Copying gs://public-datasets-deepmind-alphafold-v4/proteomes/proteome-tax_id-3702-6_v4.tar...
Copying gs://public-datasets-deepmind-alphafold-v4/proteomes/proteome-tax_id-3702-7_v4.tar...
Copying gs://public-datasets-deepmind-alphafold-v4/proteomes/proteome-tax_id-3702-8_v4.tar...
Copying gs://public-datasets-deepmind-alphafold-v4/proteomes/proteome-tax_id-3702-9_v4.tar...
- [14/14 files][ 20.5 GiB/ 20.5 GiB] 100% Done 17.8 MiB/s ETA 00:00:00
Operation completed over 14 objects/20.5 GiB.
Output file: //Volumes/SPIN_5TB/Science/ARATH_shards_foldseek/ARATH_AF2_shards_foldseekDB_2023Mar31
[=================================================================] 100.00% 14 2m 21s 40ms
Time for merging to ARATH_AF2_shards_foldseekDB_2023Mar31_ss: 0h 0m 0s 661ms
Time for merging to ARATH_AF2_shards_foldseekDB_2023Mar31_h: 0h 0m 0s 59ms
Time for merging to ARATH_AF2_shards_foldseekDB_2023Mar31_ca: 0h 0m 9s 942ms
Time for merging to ARATH_AF2_shards_foldseekDB_2023Mar31: 0h 0m 1s 316ms
Reordering by identifier
Time for merging to ARATH_AF2_shards_foldseekDB_2023Mar31_h: 0h 0m 0s 56ms
Time for merging to ARATH_AF2_shards_foldseekDB_2023Mar31_ss: 0h 0m 0s 693ms
Time for merging to ARATH_AF2_shards_foldseekDB_2023Mar31_ca: 0h 0m 9s 0ms
Time for merging to ARATH_AF2_shards_foldseekDB_2023Mar31: 0h 0m 0s 680ms
Ignore 0 out of 132903.
Too short: 0, incorrect 0.
Time for processing: 0h 3m 17s 911ms
```
Downloaded file / folder details:
```
% ls -lht
total 43056128
-rwxrwxrwx 1 aksrao staff 1.8G Mar 31 19:54 proteome-tax_id-3702-9_v4.tar
-rwxrwxrwx 1 aksrao staff 1.4G Mar 31 19:54 proteome-tax_id-3702-8_v4.tar
-rwxrwxrwx 1 aksrao staff 1.6G Mar 31 19:53 proteome-tax_id-3702-6_v4.tar
-rwxrwxrwx 1 aksrao staff 1.7G Mar 31 19:52 proteome-tax_id-3702-7_v4.tar
-rwxrwxrwx 1 aksrao staff 1.6G Mar 31 19:51 proteome-tax_id-3702-4_v4.tar
-rwxrwxrwx 1 aksrao staff 1.6G Mar 31 19:51 proteome-tax_id-3702-5_v4.tar
-rwxrwxrwx 1 aksrao staff 1.6G Mar 31 19:50 proteome-tax_id-3702-3_v4.tar
-rwxrwxrwx 1 aksrao staff 1.5G Mar 31 19:48 proteome-tax_id-3702-2_v4.tar
-rwxrwxrwx 1 aksrao staff 1.5G Mar 31 19:47 proteome-tax_id-3702-1_v4.tar
-rwxrwxrwx 1 aksrao staff 1.6G Mar 31 19:46 proteome-tax_id-3702-12_v4.tar
-rwxrwxrwx 1 aksrao staff 496M Mar 31 19:45 proteome-tax_id-3702-13_v4.tar
-rwxrwxrwx 1 aksrao staff 1.5G Mar 31 19:45 proteome-tax_id-3702-11_v4.tar
-rwxrwxrwx 1 aksrao staff 1.1G Mar 31 19:44 proteome-tax_id-3702-10_v4.tar
-rwxrwxrwx 1 aksrao staff 1.5G Mar 31 19:44 proteome-tax_id-3702-0_v4.tar
```
## Observations:
1. The pre-made DB appears to contains predicted structures for 27434 UniProt IDs - looks to me like 27434 *.pdb.gz and 27434 *.cif.gz files each, for a total of 54868 entries
2. The shard tar file download with gsutil, in contrast, has 132903 cif files, but no pdb files?
3. At https://github.com/deepmind/alphafold/blob/main/afdb/README.md, I note :
> there are 3 files per protein
4. At https://alphafold.ebi.ac.uk/download, I note:
> In the case of proteins longer than 2700 amino acids (aa), AlphaFold provides 1400aa long, overlapping fragments.
## Question 1 (test-case specific)
From UniProt download page at https://www.uniprot.org/proteomes/UP000006548, the number of ALL proteins is 39324, whereas the representative proteome (1 protein per gene) = 27500.
My understanding, from your example of Titin (Q8WZ42) at https://alphafold.ebi.ac.uk/download is that the overlapping fragments 1400aa in length, used in AF2 structure predictions, slide over by 200aa, when the full length protein is > 2700aa long.
Examining the lengths of proteins in these 2 proteome files, across ALL ARATH proteins there are 53 out of 39324 that are > 2700aa long and across representative ARATH proteins there are 22 out of 27500 that are > 2700aa long.(see 2 attached files with 3 columns each - UniProt IDs , protein lengths, and numbers of split entries)
[UP000006548_ARATH_RepProteome27500_IDs_SeqLen_AF2_2700maxLen_200SlideOverlap.txt](https://github.com/deepmind/alphafold/files/11132773/UP000006548_ARATH_RepProteome27500_IDs_SeqLen_AF2_2700maxLen_200SlideOverlap.txt)
[UP000006548_ARATH_FullProteome39324_IDs_SeqLen_AF2_2700maxLen_200SlideOverlap.txt](https://github.com/deepmind/alphafold/files/11132774/UP000006548_ARATH_FullProteome39324_IDs_SeqLen_AF2_2700maxLen_200SlideOverlap.txt)
Based on this math, the additional number of "split" entries, for representative and full proteomes should be 114 and 276 - which means an additional 114 *3 and 276*3 entries respectively, for respective totals for representative and full proteomes for _Arabidopsis thaliana_, of:
(27500-22) + 114 = 27592 entries and (27500-22)*3 + 114*3 = 82776 AF2 files (_A. thaliana_ representative proteome) This is still fewer than the
(39324-53) + 276 = 39547 entries and (39324-53)*3 + 276*3 = 118641 AF2 files (_A. thaliana_ full proteome) This is still fewer than the 132903 gif files from my shard tar file download count for _Arabidopsis thaliana_, by 132903 - 118641 = 14262. Why this discrepancy? What am I missing or doing incorrectly in my calculations?
## (general) Question 2
Using `gsutil`, is It possible at all to download ONLY 1 protein per gene, as is available for some UniProt "representative proteomes", but only for some species at UniProt? If yes, what would the download syntax be? Or would this best be implemented as a post-download parsing step ? If latter is true, please share details on best syntax / steps for such a subsetting.
## (general) Question 3
What is the correct way to check whether the number of UPID entries / number of AF2 prediction files downloaded as shard tar files are identical to downloads of pre-made species-specific AF2 database?
Thanks in advance!
Contributor guide
Assessment
This issue has not been assessed yet.