sokrypton / sokrypton/ColabFold

evoformer embeddings

Open
#586 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Jupyter Notebook
Stars
2.9k
Forks
747
PR merge metrics
No merged PRs in 30d

Description

Hi @sokrypton and thanks for another great project,

No "issue" here, rather a question since I couldn't find it in this repo or the colabdesign one ...
Is there a way to get the residue embeddings that are computed by the evoformer within the structure prediction model please?

Whether the input is a single sequence or an MSA, my understanding is that some residue embeddings are computed for the query sequence, which are then used as node features for the folding model.

I haven't found anything about how to compute these, although it could be super useful as a representation.

Thanks!

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

The issue names no file, test, or entry point. First trace how ColabFold obtains Evoformer outputs for sequence and MSA inputs, then determine whether residue embeddings can be exposed; done would be a documented, reproducible way to retrieve them or a clear explanation that this is unsupported.

Written by the indexing model from the issue text.

Assessment

Tech stack
jupyter-notebook
Domain
bioinformatics, machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.