sokrypton / sokrypton/ColabFold

High amount of clashes in multimer models

Open
#72 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Jupyter Notebook
Stars
2.9k
Forks
747
PR merge metrics
No merged PRs in 30d

Description

Hi, thanks for the update have had great use of ColabFold so far!

Using the new multimer models I very often get a high amount of clashes that can be distracting while trying to look at interactions, and hiding them can be time consuming when going through a larger amount of predictions. This only seems to happen when there is larger disordered regions. And if the disordered regions are very large the proteins can become just a blob. One less extreme example below (from your new multimer notebook);

cf_clashes_example

The disordered region after the well predicted domain is tangled and blocks the view of the interface. It also calls into question if the interface is actually possible due to the disordered region originating from that location.

This is something that I've noticed in the AlphaFold notebooks, our local AlphaFold setup and now in your new multimer notebook. Your original code for complexes however seem to always respect stereochemistry. Is this just how the new multimer models handle uncertainty or is something going wrong? And how are we supposed to interpret and work around this high amount of clashes. Hide and ignore them or just throw away the prediction as too low confidence?

Thanks for your help!

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the new multimer notebook and compare its behavior with the original complexes code and the other AlphaFold notebooks mentioned in the report. Determine whether the clashes in large disordered regions are expected model uncertainty or an implementation issue, and document how users should interpret or handle them.

Written by the indexing model from the issue text.

Assessment

Tech stack
jupyter-notebook
Domain
bioinformatics, machine-learning
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.