sokrypton / sokrypton/ColabFold

Request on how to generate extra MSA files available from public server vs locally using `colabfold_search`

Open
#703 2 comments 1 reaction 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Jupyter Notebook
Stars
2.9k
Forks
747
PR merge metrics
No merged PRs in 30d

Description

Hello ColabFold team!,

First, thank you so much for maintaining, updating and creating ColabFold! I really appreciate your team's efforts!

I noticed that runningcolabfold_batch using the public server as MSA source , I see that the output has 3 types of MSA files namely ( as shown below copied from Ref. #580 ) -- heterodimer_2.a3m , pair.a3m and uniref.a3m files.

However, when I use colabfold_search on my locally created database (database was created about ~6 months ago) I only get the MSA file heterodimer_2.a3m. Passing heterodimer_2.a3m file to colabfold_batch doesn't generate any additional MSA files as above.

So my question is --

  • are there scripts/ways to get the pair.a3m and uniref.a3m files using colabfold_search on my locally created database?
  • It will be also great if you can elaborate on what are the contents of pair.a3m and uniref.a3m files and how they affect the structure prediction accuracy.
Results from using colabfold_batch where the MSAs come from public server
.
├── cite.bibtex
├── config.json
├── log.txt
├── heterodimer_2.a3m
├── heterodimer_2_coverage.png
├── heterodimer_2.done.txt
├── heterodimer_2_env
│   ├── bfd.mgnify30.metaeuk30.smag30.a3m
│   ├── msa.sh
│   ├── out.tar.gz
│   ├── pdb70.m8
│   ├── templates_101
│   │   ├── 7x8v.cif
│   │   ├── pdb70_a3m.ffdata
│   │   ├── pdb70_a3m.ffindex
│   │   ├── pdb70_cs219.ffdata
│   │   └── pdb70_cs219.ffindex -> pdb70_a3m.ffindex
│   ├── templates_102
│   │   ├── 1t1h.cif
│   │   ├── 2c2l.cif
│   │   ├── 2c2v.cif
│   │   ├── 2f42.cif
│   │   ├── 2oxq.cif
│   │   ├── 5olm.cif
│   │   ├── 6fga.cif
│   │   ├── 6s53.cif
│   │   ├── 7bbd.cif
│   │   ├── 7c96.cif
│   │   ├── 8a58.cif
│   │   ├── pdb70_a3m.ffdata
│   │   ├── pdb70_a3m.ffindex
│   │   ├── pdb70_cs219.ffdata
│   │   └── pdb70_cs219.ffindex -> pdb70_a3m.ffindex
│   └── uniref.a3m
├── heterodimer_2_pae.png
├── heterodimer_2_pairgreedy
│   ├── out.tar.gz
│   ├── pair.a3m
│   └── pair.sh
├── heterodimer_2_plddt.png
├── heterodimer_2_predicted_aligned_error_v1.json
├── heterodimer_2_relaxed_rank_001_alphafold2_multimer_v3_model_1_seed_000.pdb
├── heterodimer_2_relaxed_rank_002_alphafold2_multimer_v3_model_3_seed_000.pdb
├── heterodimer_2_relaxed_rank_003_alphafold2_multimer_v3_model_5_seed_000.pdb
├── heterodimer_2_relaxed_rank_004_alphafold2_multimer_v3_model_2_seed_000.pdb
├── heterodimer_2_relaxed_rank_005_alphafold2_multimer_v3_model_4_seed_000.pdb
├── heterodimer_2_scores_rank_001_alphafold2_multimer_v3_model_1_seed_000.json
├── heterodimer_2_scores_rank_002_alphafold2_multimer_v3_model_3_seed_000.json
├── heterodimer_2_scores_rank_003_alphafold2_multimer_v3_model_5_seed_000.json
├── heterodimer_2_scores_rank_004_alphafold2_multimer_v3_model_2_seed_000.json
├── heterodimer_2_scores_rank_005_alphafold2_multimer_v3_model_4_seed_000.json
├── heterodimer_2_template_domain_names.json
├── heterodimer_2_unrelaxed_rank_001_alphafold2_multimer_v3_model_1_seed_000.pdb
├── heterodimer_2_unrelaxed_rank_002_alphafold2_multimer_v3_model_3_seed_000.pdb
├── heterodimer_2_unrelaxed_rank_003_alphafold2_multimer_v3_model_5_seed_000.pdb
├── heterodimer_2_unrelaxed_rank_004_alphafold2_multimer_v3_model_2_seed_000.pdb
└── heterodimer_2_unrelaxed_rank_005_alphafold2_multimer_v3_model_4_seed_000.pdb
Results from locally created MSA file heterodimer_2.a3m using colabfold_search from local database followed by passing to colabfold_batch
.
├── cite.bibtex
├── config.json
├── log.txt
├── heterodimer_2.a3m
├── heterodimer_2_coverage.png
├── heterodimer_2.done.txt
├── heterodimer_2_pae.png
├── heterodimer_2_plddt.png
├── heterodimer_2_predicted_aligned_error_v1.json
├── heterodimer_2_relaxed_rank_001_alphafold2_multimer_v3_model_1_seed_000.pdb
├── heterodimer_2_relaxed_rank_002_alphafold2_multimer_v3_model_5_seed_000.pdb
├── heterodimer_2_relaxed_rank_003_alphafold2_multimer_v3_model_3_seed_000.pdb
├── heterodimer_2_relaxed_rank_004_alphafold2_multimer_v3_model_2_seed_000.pdb
├── heterodimer_2_relaxed_rank_005_alphafold2_multimer_v3_model_4_seed_000.pdb
├── heterodimer_2_scores_rank_001_alphafold2_multimer_v3_model_1_seed_000.json
├── heterodimer_2_scores_rank_002_alphafold2_multimer_v3_model_5_seed_000.json
├── heterodimer_2_scores_rank_003_alphafold2_multimer_v3_model_3_seed_000.json
├── heterodimer_2_scores_rank_004_alphafold2_multimer_v3_model_2_seed_000.json
├── heterodimer_2_scores_rank_005_alphafold2_multimer_v3_model_4_seed_000.json
├── heterodimer_2_template_domain_names.json
├── heterodimer_2_unrelaxed_rank_001_alphafold2_multimer_v3_model_1_seed_000.pdb
├── heterodimer_2_unrelaxed_rank_002_alphafold2_multimer_v3_model_5_seed_000.pdb
├── heterodimer_2_unrelaxed_rank_003_alphafold2_multimer_v3_model_3_seed_000.pdb
├── heterodimer_2_unrelaxed_rank_004_alphafold2_multimer_v3_model_2_seed_000.pdb
├── heterodimer_2_unrelaxed_rank_005_alphafold2_multimer_v3_model_4_seed_000.pdb
└── templates
    ├── 1t1h.cif
    ├── 2c2l.cif
    ├── 2c2v.cif
    ├── 2f42.cif
    ├── 2oxq.cif
    ├── 5olm.cif
    ├── 6fga.cif
    ├── 6s53.cif
    ├── 7bbd.cif
    ├── 7c96.cif
    ├── 7x8v.cif
    ├── pdb70_a3m.ffdata
    ├── pdb70_a3m.ffindex
    ├── pdb70_cs219.ffdata
    └── pdb70_cs219.ffindex

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by tracing the colabfold_search and colabfold_batch entry points and comparing the listed public-server and local output trees. Check how heterodimer_2.a3m, pair.a3m, and uniref.a3m are produced; done means documenting whether local generation supports both files and explaining their roles.

Written by the indexing model from the issue text.

Assessment

Domain
bioinformatics
Issue type
Documentation
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.